Documentation
¶
Overview ¶
Package textchunk cuts a text longer than an embedding provider's input into pieces that each fit, so the whole text reaches the model across several calls rather than being trimmed (#1242, #2027). The caller says where its kind of text breaks naturally (a markdown heading, a Starlark definition); the rest is blank lines, then words, then rune boundaries.
Index ¶
Constants ¶
const MinViableChunkBytes = 64
MinViableChunkBytes is the smallest input budget worth splitting to. Below it, splitting stops being meaningful (a chunk would hold a few words) and a hard cut can no longer be guaranteed to make progress: the budget must exceed the widest UTF-8 rune, or a rune-boundary cut could produce an empty piece and the split would not terminate. A caller handed a budget this small embeds its text whole and the provider's own cap applies.
Variables ¶
This section is empty.
Functions ¶
func SplitText ¶
SplitText cuts text into pieces of at most budget bytes each, so a text longer than the provider's input reaches the model whole, across several calls, rather than being trimmed (#1242, #2027). sections cuts the text at the boundaries its kind has (a markdown heading, a Starlark definition); each section that does not fit is cut at blank lines, then at a word on a rune boundary, and consecutive pieces are packed back together while they fit. sections must be lossless: the pieces it returns concatenate back to the text. Whitespace-only pieces are dropped, since they carry nothing a vector could match.
func TruncateOnRune ¶
TruncateOnRune returns the longest prefix of s within maxBytes that ends on a UTF-8 rune boundary. A non-positive maxBytes returns s unchanged.
Types ¶
This section is empty.