Documentation
¶
Overview ¶
Package llmlab is LLMLab, an OpenAI-compatible chat completions API that streams made-up tokens over server-sent events. It is the reference app for the llm-apps pack.
Planted bottlenecks (see README.md), each switched off by a fix flag:
- slots: only four generations run at once; the rest queue before their first token, so time to first token grows with concurrency (fix "batch" allows 256).
- prefill: prompt processing holds one global lock, so a long prompt delays the first token of every other request (fix "chunked").
- usage: usage accounting re-tokenises the whole reply after every token, so CPU per stream grows with the square of its length (fix "usage").
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
Types ¶
This section is empty.
Click to show internal directories.
Click to hide internal directories.