Inference Splits Into Specialized Markets as Local and Voice AI Expand
Inference was described as the largest market in software, replacing databases, with segmentation by speed, cost and workload type. An investment in Sail targets cases where users can accept a five-to-ten-minute wait in exchange for substantially lower inference costs; many business applications were characterized as more tolerant of delay than consumer AI.
The same investor cited voice inference, robotics and local AI as distinct areas of focus. Ollama was described as initially being used by about 10 million people for local AI, while a Stanford study reportedly released three or four weeks earlier found that 90% of white-collar AI use cases could be handled on a MacBook. The excerpt does not provide the study’s methodology or exact date. More demanding workloads can be routed to the cloud.
Voice AI can run on a device, at the edge or in a data center, but its architecture includes telecommunications, call digitization and model processing, making latency a key concern. KV-cache warming can prepare systems for time-zone-driven demand. Voice inference is expected to begin mainly in centralized data centers and likely move toward the edge over time.
