Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Summary
Tencent WorkBuddy Bench is a new multi-domain evaluation suite for coding agents, featuring a contamination-resistant task construction methodology. It utilizes a unified framework to generate distribution-informed tasks across four key work domains: Code, Web, Office, and Security. Tasks are reverse-engineered from real commits, pull requests, or business scenarios and rewritten as colloquial, role-played requests, preventing web-search discoverability of original issues. The dataset, including task directories and evaluation harnesses, is openly released, with contamination resistance maintained through its unique construction and versioning. The suite includes four distinct subsets—repository-level engineering, front-end development, office/business workflows, and red-/blue-team security—each with specific verification styles. Tasks run under a uniform protocol on CodeBuddy Code and Claude Code agent harnesses, ensuring reproducibility and auditability. A cross-model leaderboard is reported, though subset scores are not directly comparable. The benchmark was published on 2026-07-23.
Key takeaway
For AI Engineers evaluating coding agents, the Tencent WorkBuddy Bench offers a robust, contamination-resistant resource. You should consider integrating this benchmark to assess your agents' performance across diverse real-world scenarios, including Code, Web, Office, and Security domains. Its open release and uniform protocol ensure reproducibility and auditability, allowing you to confidently compare models and identify specific areas for improvement in agent capabilities.
Key insights
The Tencent WorkBuddy Bench offers a contamination-resistant, multi-domain benchmark for coding agents, ensuring robust and reproducible evaluation.
Principles
- Task construction should prevent data contamination.
- Openly releasing benchmarks enhances auditability.
- Multi-domain evaluation reveals agent versatility.
Method
Tasks are reverse-engineered from real commits/PRs/scenarios, then rewritten as colloquial, role-played requests to avoid web-search discoverability. This ensures contamination resistance.
In practice
- Evaluate coding agents across diverse domains.
- Use reverse-engineered tasks for robust testing.
- Inspect benchmark content for auditability.
Topics
- Coding Agents
- Benchmark Datasets
- Contamination Resistance
- Multi-Domain Evaluation
- Tencent WorkBuddy Bench
- Software Engineering
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, AI Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.