Sorry that I write in this topic, couldn't find another way to reach out to you.
Hi! I’m maintaining the Explicit Edit Benchmark, a small public benchmark for measuring how much agent tooling and harness design affect a model’s ability to make precise text edits.
Text editing is one of the most common operations in coding-agent workflows. In my experience, it is also one of the operations agents struggle with the most. Encodings, text blocks, line endings, invisible characters, unicode, and file formats with strict formatting or structural requirements create a huge number of variations. I have seen an agent fail to edit a single line because it could not reproduce an invisible character in that line. There are many cases like this.
The benchmark currently contains 226 deterministic tasks across different file formats and edit types: replacements, insertions, deletions, moves, copies, unicode cases, large files, and other edge cases. The tasks are intentionally simple, so they can be run even with low reasoning and keep the focus on the model and tooling combination rather than deep problem solving.
I’ve been running as many agents, models, configurations, and Pi editing extensions as I can and publishing every accepted observation in the public dataset. There are too many combinations for one person to cover, and model behavior is stochastic, so a single run cannot give us a reliable picture. The benchmark was designed from the start as a public, community-driven dataset that anyone interested and able to run it can contribute to. We can get reliable information only by collecting enough observations across different setups.
I'm reaching out to other developers of harnesses in case they are willing to participate in a small competition and, potentially, improve their tooling as the result. Once you have an entry, you can get a nice badge like
.
Explorer: https://huggingface.co/spaces/alexshpunt/benchmark-explorer
Dataset: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark
Benchmark source: https://github.com/alexshpunt/explicit-edit-benchmark
If you’re interested, you’re very welcome to try it out, otherwise no actions needed!
Sorry that I write in this topic, couldn't find another way to reach out to you.
Hi! I’m maintaining the Explicit Edit Benchmark, a small public benchmark for measuring how much agent tooling and harness design affect a model’s ability to make precise text edits.
Text editing is one of the most common operations in coding-agent workflows. In my experience, it is also one of the operations agents struggle with the most. Encodings, text blocks, line endings, invisible characters, unicode, and file formats with strict formatting or structural requirements create a huge number of variations. I have seen an agent fail to edit a single line because it could not reproduce an invisible character in that line. There are many cases like this.
The benchmark currently contains 226 deterministic tasks across different file formats and edit types: replacements, insertions, deletions, moves, copies, unicode cases, large files, and other edge cases. The tasks are intentionally simple, so they can be run even with low reasoning and keep the focus on the model and tooling combination rather than deep problem solving.
I’ve been running as many agents, models, configurations, and Pi editing extensions as I can and publishing every accepted observation in the public dataset. There are too many combinations for one person to cover, and model behavior is stochastic, so a single run cannot give us a reliable picture. The benchmark was designed from the start as a public, community-driven dataset that anyone interested and able to run it can contribute to. We can get reliable information only by collecting enough observations across different setups.
I'm reaching out to other developers of harnesses in case they are willing to participate in a small competition and, potentially, improve their tooling as the result. Once you have an entry, you can get a nice badge like
.
Explorer: https://huggingface.co/spaces/alexshpunt/benchmark-explorer
Dataset: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark
Benchmark source: https://github.com/alexshpunt/explicit-edit-benchmark
If you’re interested, you’re very welcome to try it out, otherwise no actions needed!