Tenvor documentation

Models and limits guide

Use the current catalog for model choices and numeric limits. This guide explains the constraints without presenting a static capacity or price snapshot as live data.

Content checked . Live limits and prices can change.

Choose a current model ID

Models you can address by public tag are shown in the live catalog. External clients can discover authorized IDs through GET /v1/models. Use the exact ID returned; do not derive a model alias from an advertised family name or a historical release note.

A model listing explains a purchasable interface; it does not promise an idle worker, a specific GPU or a completion time. If the catalog is unavailable, wait for it before choosing current limits or prices.

Three separate limits

  • max_context_tokens is the model's published context constraint. Input history and requested output must fit the applicable admission policy.
  • max_input_bytes is an aggregate input-byte budget, not a character or input-token allowance. Framing, tool schemas, call arguments, results and conversation history consume budget; multibyte text can use more than one byte per character.
  • max_output_tokens caps the requested completion. A lower value is a ceiling; it does not force that many tokens or guarantee a complete long answer.

The catalog is published at GET /api/catalog/public. Admission enforces the current request contract. If a request is rejected for context, reduce history, attachments or tool definitions and choose a supported output budget rather than inventing a larger limit.

Check prices separately from capacity

Use the pricing calculator for current request minimums and token rates. An estimate does not reserve funds or compute; admission and final settlement are separate steps.

Cross-request prefix reuse requires an explicit client opt-in and an accepted server lease. The OpenCode setup script installs the required cache plugin; ordinary requests without opt-in run cold. The server may decline retention or fall back to cold work. Reuse does not currently imply a billing discount, Fast tariff or latency SLA.

Read the API contract · Understand reservations