Jina - Segment text action ​

The Segment text action splits text into semantic chunks with Jina Segmenter in Jina.

Input ​

Input fieldDescription
ContentEnter the text to segment. The maximum is 64,000 characters.
Return chunksOptional. Select Yes or No to return the text chunks.
Max chunk lengthOptional. Enter the maximum length of each chunk.
TokenizerOptional. Select the tokenizer to use for counting tokens, such as cl100k_base or gpt2.
Return tokensOptional. Select Yes or No to return the tokenized input.
Head (tokens)Optional. Enter a number to return only the first N tokens. Don't use this setting in combination with Tail (tokens).
Tail (tokens)Optional. Enter a number to return only the last N tokens. Don't use this setting in combination with Head (tokens).

Output ​

Output fieldDescription
Num tokensThe number of tokens in the input text.
TokenizerThe tokenizer used.
UsageToken usage.
Tokens (Usage)Token usage, matching Num tokens.
Num chunksThe number of chunks returned.
Chunk positionsAn array of start and end character offsets for each chunk.
ChunksAn array of the text chunks. The Chunks field is only available when Return chunks is Yes.
TokensAn array of the tokenized input. The Tokens field is only available when Return tokens is Yes.

Last updated: