Spec-Driven Delivery · Claude Code
A toolkit that turns Claude Code into a spec-driven delivery line — requests are grilled into specifications, plans are challenged until no gaps remain, and code ships only after gated review and a verification loop.
Claude is the primary engine and drives every phase. Codex or Kimi is an optional second engine that you select and configure through /m:setup.
/m is a set of twenty Claude Code commands and skills that run software delivery as one disciplined pipeline instead of a chat. The work moves through gated phases, each one marking itself done before the next is entered, and the project's own memory is read at every step. Three parts make up the system:
Pipeline
Gated delivery
One command — /m:develop — drives every change through five phases: refine, plan, implement, review, iterate. Each runs as a discrete skill behind a marker-file gate.
Support · memory
Project state
Utility stages build and feed the per-repo .m/ memory — index, tasks, progress, gaps, research — the shared state every pipeline stage reads.
Expert modes
Review lenses
Specialist review modes clip onto the work — some run by hand, others load themselves when matching files are edited — to harden security, Go, React, and business logic.
/m:develop orchestrates the five phases below. A hook backs the work: while a run is live, it blocks edits to your project files until the current phase has been entered through its skill. The order itself — refine before plan before implement — is the protocol /m:develop steps through, checking each phase's marker before the next.
Claude, Codex, and Kimi are engines, not models — each runs its own underlying models. The pipeline runs on Claude by default and works fully Claude-only. Codex or Kimi is a dual-engine add-on: when one is selected, a second engine independently checks the plan, the research, and the review, and disagreements are surfaced side-by-side rather than silently merged.
Claude Code is the surface and Claude is the engine driving every phase — refine, plan, implement, review, iterate. Its models are used across stages by weight of task: Opus 5.5 for refine, plan, implement, and review, Sonnet 5 for iterate and status, Haiku 4.5 for help, and Fable 5.1 for analyze. Nothing else is required: with no second engine, every stage runs Claude-only, minus the second opinion.
OpenAI's Codex CLI or the Kimi Code CLI, used as an optional second engine for /m:plan, /m:research, and /m:review. When selected it runs automatically on every applicable pass and is token-metered per run, so its cost stays capped.
You pick the provider — codex, kimi, or none — and its model (gpt-6-astra, gpt-6-sol, …), reasoning effort, and fast mode with the /m:setup wizard. Settings live in each repo's .m/pipeline.yml; the default is provider: none, which is Claude-only.
Getting running is short. Claude Code is all you need for the full pipeline; a second engine is a deliberate, optional second step.
Everything in /m runs on top of Claude Code. Once it is installed and authenticated, the pipeline is usable end-to-end on Claude alone.
The repository is a Claude Code plugin marketplace. Run /plugin marketplace add milorad-teodorovic/m-pipeline, then /plugin install m@m-pipeline. With the plugin installed, /m:develop and the individual phases are available in any project.
Want the dual-engine second opinion? Run the guided setup wizard. It checks the CLI and login of each provider, then walks you through the choices and writes them for you:
Codex needs its own CLI and a ChatGPT login; Kimi needs the Kimi Code CLI and user-level deny rules. /m:setup tells you exactly what is missing and how to fix it. Read-only check first? Run /m:setup --check.
Each repo's .m/pipeline.yml holds the second_engine: block — provider, fast_mode, model, reasoning_effort, and the token budget. Absent file means built-in defaults apply; /m:setup can write it for you.
The repository ships 37 eval cases that run each stage through claude plugin eval in a scaffolded workspace. Each trial is PASS, FAIL, or INVALID; an invalid trial never counts as a pass.
Fixture preflight, grader regression tests, and hook tests.
One trial per selected case, with a cost cap.
Runs the same suite against another plugin checkout. Each report records the models, the CLI version, and hashes of the plugin, fixtures, and graders, so an incompatible comparison is rejected.
Two companion pages go deeper — the full interactive pipeline schematic, and the statusline that rides along in your terminal.
An engineering working-drawing of all twenty stages — five gated phases, support and expert lanes, dual-engine checks, and per-repo memory. Select any component to read its live specification.
Open the schematic → Terminal companion m-statuslineA single-file Claude Code statusline — model, context burn, and pipeline state at a glance, with the pixel-burn duck brand. Drops into your Claude Code config.
Open the statusline page →