Try a skill.
Pick any of the six below and you land inside the app with the whole workflow already loaded, briefed, and waiting for your first command.
From command to production.
Nothing below is mocked for the page, because the commits are actual production deploys, the plan is a live run mid-flight, and the scaffold is exactly what the CLI hands you when a new skill is born.
What shipped, when.
Every entry is a real production deploy. Subjects come straight from the commits, no rewrites.
View full changelogPlan, run, observe, iterate.
Ask in plain English. Ultron drafts a plan, runs the skills, and pauses at every decision that leaves the building.
Workflow runtimeRun a skill.
Trigger any of the 57 skills from chat with a slash command or plain language.
View slash commandsBuild a skill.
Scaffold a skill with a manifest, a tool allow list, and a model tier. Loaded on next start.
Skill anatomySkills, techniques, and field notes.
Everything deeper lives one click from here, from the full skill catalog to the hand-crafted workflows and the teardowns we wrote after running real campaigns.
Every frontier model, one runtime.
Whichever lab ships the best frontier model this month, it is already wired into the same runtime, so you can pin one per skill or let the router decide on every turn.
How they stack up.
Every model runs the same benchmark, where higher quality and throughput win the column and lower violation rate and cost win theirs.
- 01
gpt-5.50.82
- 02
claude-opus-4-60.82
- 03
gpt-5.40.81
- 04
claude-opus-4-50.81
- 05
Kimi-K2.60.79
- 06
DeepSeek-V4-Pro0.78
Indicative benchmark figures. Accent marks the best result in each column.
Six systems that make it run.
Every turn you run leans on the same quiet machinery underneath, six subsystems passing work down one line from the skill you invoke to the background job that quietly finishes it.