Model Configuration in Zoho Zia Agents determines which large language model (LLM) your agents use when processing tasks and generating responses. You can choose from Zoho-hosted models via the Zoho Key System (ZKS) or supply your own API credentials through Bring Your Own Key (BYOK). This article explains both options, the models available, and how pricing works.
ZKS gives you access to a range of powerful LLMs without needing to set up your own API keys. Zoho manages the connection to supported providers and handles billing on your behalf. Because these models run within Zoho's infrastructure, your data remains inside the Zoho environment, which may be important for organisations with data residency requirements.
The following models are currently available through ZKS:
BYOK allows you to connect your existing API credentials from an LLM vendor directly to Zia Agents. When you use BYOK, there is no cost on the Zoho platform side; the LLM provider bills you directly according to their own pricing. This option suits organisations that already have preferred vendor agreements or need to use a specific model not covered by ZKS.
Supported BYOK vendors are OpenAI, Anthropic, Google Gemini, DeepSeek, and Cohere.
The right choice depends on your organisation's circumstances:
ZKS models fall into two pricing tiers. Each tier includes a monthly free token allowance, with usage above that threshold consuming AI credits.
| Tier | Models Included | Free Monthly Tokens | Overage Rate |
|---|---|---|---|
| Standard | Qwen 14B, Qwen 3.5 35B MoE, GLM 4.7 Flash | 30 million | $1.00 per 1M tokens |
| Pro | GLM 5 | 20 million | $3.00 per 1M tokens |
The credit exchange rate is 1 USD = 1,000 AI Credits.
Older models are periodically retired as vendors update their offerings. Deprecated models include certain OpenAI GPT-4 preview variants, selected Anthropic Claude versions, and Zoho's Qwen 30B MoE. If you are using a model that is approaching retirement, Zoho will notify you in advance. Ensure you update your agent's model configuration before a model is fully withdrawn to avoid disruption.