TBPN

← Full issue

September 4, 2026

Expert testing proposed as a complement to AI benchmarks

The benchmark era may be becoming less useful for evaluating AI model launches because published tests can be optimized quickly, reaching very high scores before benchmark discussions lose much of their meaning.

A practical alternative or complement is to test a new model on a domain the evaluator understands intimately and assess whether its responses are genuinely impressive. Domain-specific checks—such as evaluating a model on horses for someone with deep knowledge of horses—could expose weaknesses that benchmark scores miss.

Privacy ·