Coding agent
/weco improve validation accuracy
The search is a tree. Open any node to see its code, terminal output, and metric.
Try a paper or your own idea; it becomes a new branch of the search.
GPT, Claude, Gemini, Kimi, GLM, MiniMax, with bring your own key.
AUC 0.91 → 0.93
for end-to-end transaction-fraud ML pipelines
+32.4% accuracy
for vision-language chart-to-CSV extraction
+47.7% speedup
for causal self-attention in GPT-style LLMs
−91.3% RMSLE
for molecular property prediction in materials science
Type /weco in Claude Code, Codex, or Cursor. The agent sends Weco your task; Weco iteratively writes the next code version and runs your eval.
/weco improve validation accuracy
So amazing to see something built by this team that's substantially underpinning and influencing OpenAI's agentic roadmap.

Weco consistently surfaced improvements across my entire ML pipeline that I would never have discovered manually. It doesn't just accelerate development - it enables a fundamentally different way of building intelligent systems.

OpenAI is nothing without AIDE. Really cool to see @WecoAI's agent framework give o1 such a huge uplift.

Weco is how we go from 90% to 99%. We use it as a second phase after training - and the improvements it finds have directly helped us close deals.


Setting up Weco was incredibly easy; it instantly understood what to do, and I barely had to touch it.

We used Weco to optimize our voice AI prompts - accuracy jumped from 82% to 96%, with entire failure categories going from 0% to 100% in just a few iterations.

We used Weco on a bioacoustics research problem where we had to design new algorithms from scratch. It replaced weeks of manual heuristic work.

Try it on a real run
Bring your own keys
Managed inference at scale
Connect your coding agent to the dashboard and turn any eval into a self-improving loop.