BaseLabs advocates evaluating AI models through real-world economic use
BaseLabs argues that conventional AI benchmarks often target narrow failure modes. ArcAGI is cited as an example of a specialized area where models historically performed poorly; once a benchmark is public, it can also become a target for reinforcement learning and lose some independence.
The company says there is no set of 10 benchmarks that can fully capture a model’s strengths and uneven frontier. In its view, the strongest signal comes from aggregating evaluations built by companies using models for real, paid tasks. BaseLabs says its thousands of users and scaled RL environments can provide that signal, while acknowledging there is no universal measure of model usefulness.
