Decagon routes agent tasks to smaller models to cut cost and voice latency
Decagon says that after an AI-agent product is working, many model calls do not require a full frontier model. It uses smaller—and specifically open-source—models for internal tasks such as choosing a topic and detecting hallucinations, while retaining frontier models for new products and capabilities.
The company says smaller models have performed at the same level for these use cases after fine-tuning or post-training. For voice agents, Decagon considers latency at least as important as cost and says most current gains come from reducing model size, rather than from chipset changes or edge deployment. It suggests that edge computing could become more relevant once models are further optimized.
