A new study identifies 'token maxing' as a key driver of rising AI spend, where organizations increase reasoning depth and context size faster than task value. The research isolates the orchestration layer as the decisive lever for control, testing six foundation models while swapping only the harness design. Results suggest that better context assembly, tool exposure, and turn sequencing can significantly reduce token consumption without compromising capability.
- Token efficiency depends more on orchestration logic than the underlying foundation model chosen.
- Stop scaling token limits blindly; optimize context assembly and tool delegation first.
- Implementing stricter governance and observability in the harness curbs runaway token usage.
- Benchmarking should fix the harness to isolate model performance from orchestration waste.