Admin Release
v0.261.001
Per-Model Output Token Ceilings
Admins can set optional output-token ceilings per global model endpoint instead of relying on one tenant-wide response limit.
Visual tour
Screenshots
These screenshots come from the same app-side Latest Features gallery used in the product experience.
Current release version for Per-Model Output Token Ceilings: 0.261.001
Each model in the global multi-endpoint GPT configuration can now carry its own output-token ceiling. The chat path applies the correct backend token parameter for GPT-5 and o-series models as well as other OpenAI-compatible providers.
Why It Matters
This matters because administrators can balance cost, latency, and answer depth independently for each deployed model.
How to Try It
- Open Admin Settings > AI Models and review each global GPT endpoint.
- Set an output-token ceiling for high-cost or latency-sensitive models that need tighter limits.
- Leave the ceiling empty for models that should keep provider or application defaults.
- Test representative prompts after changing limits to confirm responses remain useful for end users.
Where to Find It
- Open Model Endpoints — Set per-model output token ceilings.