Research1 min read
Harness Performance: A Contamination-Controlled Study
A study comparing Claude-agent-sdk and deepagents across Claude-opus-4-8 and openai-codex with gpt-5.5, gemini-3.5-flash and deepseek-v3.2 found no significant advantage for either vendor-native harness. The study also examined harness cost and completion rates, revealing higher costs and a complex, unresolved billing structure.
From arXiv cs.AI


