# What Is Agentic Engineering? Definition, Practice and the Agentic Software Factory

> What is agentic engineering? Definition, origin, guardrails, autonomy stages and the agentic software factory for AI coding in the enterprise.

- URL: https://seiler.it/themen/agentic-engineering/en.html
- Author: Dr. Sven Seiler (https://seiler.it/)
- Type: Pillar / Topic Hub
- Language: en
- Updated: 2026-09-27

TL;DR

- Agentic Engineering is the discipline by which teams deploy Agentic AI across the software development lifecycle in a production-ready way.

- It addresses the core problem of the Agentic era: tooling has outpaced methodology. Gartner predicts (June 2025)[^1] that over 40% of Agentic AI projects will be canceled by the end of 2027.

- Three building blocks: specification as runtime component, six guardrails (CI, static types, linting, architectural tests, behavioural tests, MCP-driven code quality), and new team roles (context engineer, AI ethics advisor).

- Unlike vibe coding: every output is verifiable, every decision documented, every architecture controlled.

- My take: small teams with well-instrumented agents can keep pace with much larger teams — the bottleneck shifts from coding to review, testing, and QA.

## What is Agentic Engineering?

In short
Agentic Engineering is a structured, verifiable approach by which development teams deploy Agentic AI (autonomous AI agents) across the full software lifecycle — without losing control of architecture, quality, and maintainability.

The term deliberately distances itself from “vibe coding,” the unstructured use of AI coding tools where prompt outputs land directly in production.

Where vibe coding works astonishingly well for prototypes, it falls apart in production at the first schema change. Agentic Engineering is the answer to that gap: an engineering framework that treats AI output like any other code — with specification, verification, code review, and quality gates.

**Where the term comes from:** The term was popularised in early 2026 by Andrej Karpathy, co-founder of OpenAI and former head of AI at Tesla. On 4 February 2026, one year after he had coined “vibe coding”, he named “agentic engineering” his favourite[^2] for what comes next: programming via AI agents as the default, but with more oversight and without compromising on quality. My work is about how to turn that into a robust enterprise engineering discipline.

## Agentic coding, vibe coding, agentic engineering: the difference

The three terms are often mixed up, but they describe different things:

- Agentic coding is the technique: an AI agent does not just suggest code, it edits files, runs commands and uses tools itself, in the IDE or the terminal. It corresponds to rungs 3 to 5 of the agentic coding ladder below.

- Vibe coding is the uncontrolled use: prompt, accept the output, hope. Surprisingly good for prototypes, risky for production software.

- Agentic engineering is the method: specification, guardrails and clear responsibilities between human and agent, so that agentic coding becomes production-ready in the enterprise.

## Why now: the methodology–tooling gap

2025/2026 is the Agentic era: 85% of developers regularly use AI tools[^3], Cursor was valued at $29.3B in November 2025[^4], and Claude Code reached $1B in run-rate revenue six months after launch[^5]. But the tagline shift from “autocomplete” to “autonomous” happens faster than engineering practices adapt.

The flip side: Gartner predicts (June 2025)[^1] that over 40% of Agentic AI projects will be canceled by the end of 2027 — due to escalating costs, unclear business value, or inadequate risk controls. My reading: the models aren’t the problem; the gap between “technically possible” and “production-ready” is. That is exactly where Agentic Engineering steps in.

> Coding is cheap. Software is not. The gap between those two is where the real work happens — and where the real value will live for the next decade.

## The five-rung ladder of agentic coding

A mental model for where a team stands on the curve:

1. Chat — ask the model, copy the answer.

1. Mid-loop generation — AI generates code chunks, human stitches.

1. In-the-loop agentic — AI operates inside the IDE with access to files, terminal, tools.

1. On-the-loop agentic — AI works with reduced supervision. Humans set goals, agent executes, humans review.

1. Multi-agent coding — multiple specialised agents collaborate in parallel.

Most developers sit between rung 2 and 3. The interesting work happens at rungs 4 and 5. **Key insight:** Each higher rung means less human review per line of code — which only works if the surrounding system (specifications, guardrails, verification layers) gets correspondingly stronger.

## From developer to organisation: autonomy only as far as verification carries

The coding ladder describes how an individual developer works with agents. For a company that is not enough. There the question is: what may an agent do on its own in which phase, and how do we notice when it gets it wrong? For this I developed a model at Storm Reply together with colleagues. It is built from our own development work with agents and against the published frameworks of DORA[^6], DX[^7], Augment Code[^8] and Microsoft[^9].

> Autonomy may only go as far as verification carries. Degree of autonomy is not maturity – not every phase should be at stage 5.

Each stage therefore has two sides: what the agent does on its own (autonomy) and who or what secures the result (verification).

1. Assisted – the agent suggests, the human reads along.

1. Delegated – the agent completes whole work packages, the human checks every line. The most dangerous stage, because it feels like progress: the errors are now conceptual, the code looks right.

1. Supervised – a technically enforced gate checks, the human decides.

1. Autonomous – the gate checks and decides, the human only sees defined risk classes.

1. Self-directing – the loop closes from operations data: production errors become tests, telemetry becomes tickets.

**The rule:** autonomy and verification are rated separately per phase, and the minimum counts. A team whose agent takes tickets all the way to a pull request (autonomy 3) but where a human still reads every line (verification 2) is at stage 2. If autonomy runs ahead of verification, that is not progress but a risk finding.

### The map: eight phases

| Phase | Guiding question | Target stage |
| --- | --- | --- |
| Requirements | Where do requirements come from – and who writes them down? | 4 |
| Design & Architecture | Who decides architecture, security and performance? | 3 |
| Implementation | How much code is written without a human at the keyboard? | 4 |
| Test & Evaluation | How is quality measured when code gets cheap? | 4 |
| Review & Acceptance | Where does the human stay in the decision path? | 3 |
| Deploy & Operations | How does the loop close from operations back to requirements? | 4 |
| People & Roles | Who works this way – and who grants autonomy? | 4 |
| Value & Metrics | What shows that more autonomy is better? | 4 |

_The eight phases of the map with their guiding questions and the recommended target stage. Stage 5 is the default target in no phase._

### The five foundations

Beneath the phases lie five foundations. They get a traffic light rather than a stage, because they are prerequisites: **Context & Knowledge**, **Tools & Sandboxes**, **Guardrails, Hooks & Observability**, **Security, Compliance & IP** and **Cost & Token Economics**. A red foundation caps every phase that depends on it – no matter how well the agent codes.

The coding ladder and the map do not contradict each other: the ladder describes how far a developer goes, the map how far the organisation can carry it.

## The Agentic Software Factory: seven layers

The term “software factory” dates back to the 1970s and meant the standardised, repeatable production of software. In 2026 it returned as the “agentic software factory”, among others at BCG Platinion[^10] and Red Hat[^11]. My definition: **the agentic software factory is not a product you buy but the environment in which agents are allowed to work safely** – with ordered context, limited permissions, enforced checks and visible costs. With it, you can raise autonomy phase by phase. Without it, you have fast agents in a workshop without safety goggles.

1. Context – the factory's memory: versioned project rules in the repository, reusable skills and live access via MCP to tickets, docs and operations data.

1. Workflow – tickets as the agents' queue: labels trigger specification, implementation or escalation.

1. Agent runtime – every agent works in an isolated environment with its own non-human identity. Beyond a few teams you need a central agent hub. Permissions written in a prompt are not permissions.

1. Quality and gates – tests check the code, evals check the agent run. Risk classes define which change a human always sees.

1. Operations and learning loop – deploy rights granted in steps, canary and rollback as routine, incidents turn into new tickets and tests.

1. Cost – model routing, iteration limits and hard caps. The result is a number hardly anyone has today: cost per ticket.

1. People and roles – you can buy tools, not ways of working. The work shifts from writing to specifying and checking.

The research shows why this is needed. According to DORA, 90% of respondents use AI at work, and AI adoption goes along with higher throughput but still lower delivery stability[^6]. Bain puts the share of writing and testing code in the time from idea to launch at 25 to 35%[^12] – speed up only that part and the bottleneck just moves. In a randomised study by METR, experienced open-source developers took 19% longer with AI tools, yet believed they had been faster[^13]. And according to Veracode, 45% of code samples generated by language models contained OWASP Top 10 vulnerabilities[^14].

The factory does not replace architects, reviewers or acceptance. It moves their work to where it has effect: to the plan instead of the diff, to the rule instead of the single case. And it is no big bang – each layer can be built on its own. In my talks “[How to build an Agentic Software Factory](https://www.meetup.com/agentic-shift/events/316695452/)” and “[Agentic Software Factory on AWS](https://www.meetup.com/dortmund-aws-user-group/events/316695439/)” (November 2026, both with Henning Teek) we show this using our agent hub as the example.

## Specification as runtime component

In classic software engineering, requirements are a documentation artefact: nobody reads them regularly, they drift from the code. In Agentic Engineering, requirements become a **runtime component**: system prompts, skill definitions, tool descriptions, acceptance criteria are what the agent reads *every single time* it makes the next decision.

Consequence: a vague spec means a guessing agent. An ambiguous tool description is a live bug. A missing edge case is a production incident waiting to happen. This also changes the economics: time invested in precise specs compounds — every future agent invocation benefits.

## The six guardrails

Specifications tell the agent *what* to build. Guardrails tell the system *what to reject* when the agent gets it wrong. Both are necessary:

1. Continuous integration with short-lived branches — agent-generated code volume breaks classic Git workflows. Branches live for hours, not days.

1. Statically typed languages — the compiler is the cheapest, fastest, most reliable feedback loop. Domain types (PersonId instead of string) eliminate argument-swap bugs.

1. Deterministic linting — Prettier, ESLint, CSharpier. Never let the AI format code.

1. Architectural unit tests — ArchUnit and friends enforce design constraints programmatically. The agent doesn’t need to remember the architecture; the build fails when it’s violated.

1. Behavioural tests, not coverage tests — 100% coverage targets lead to AI slop tests. What hurts when it breaks — that’s what gets tested.

1. Code quality tools with MCP — SonarQube, CodeScene over the Model Context Protocol. Quality reports feed back directly to the agent — which then refactors autonomously.

## What this means for teams

The Mythical Man-Month logic starts to wobble. In my experience, small teams with well-instrumented agents can keep pace with much larger teams — not because they’re heroic, but because coordination overhead shrinks. Lines-of-code is no longer the bottleneck; review, testing, and QA become the constraint.

The funnel changes too. As an illustration (made-up numbers): previously, 500 user problems got narrowed down to 15 prioritised and 5 shipped — because coding was expensive. Today, every specified use case can get written. The filter moves forward to spec, backward to review/QA. That’s where the most valuable human work of the coming years sits.

New roles emerge: **context engineer**, **AI ethics advisor**, **AI product owner**. Classic junior coding loses weight — in the LeadDev AI Impact Report 2025[^15], 54% of surveyed members of the LeadDev community expect AI coding tools to reduce junior hiring in the long term.

## From principle to enterprise delivery: Silicon Shoring

In the enterprise, Agentic Engineering needs a robust delivery model. **Silicon Shoring** is the **Reply Group**’s AI-powered software delivery model (introduced 2025): generative AI agents across the entire software development life cycle — from requirements and code generation through testing, deployment, operations, and monitoring — blended with human expertise. It operationalizes exactly the principles on this page: specification as a runtime component, end-to-end guardrails, verifiable outputs.

At Storm Reply, Silicon Shoring is delivered **on and with AWS** — using best-of-breed agentic tooling (including Claude Code and MCP), enterprise-grade and sovereignty-aware. Two engagement models: *In-House* (the agentic system runs inside the client’s environment, integrated with their data, tools, and infrastructure) and *Managed* (a fully operated, AI-native software factory).

## Frequently asked questions about Agentic Engineering

What is an agentic software factory?

An agentic software factory is not a product but the environment in which AI agents are allowed to build software safely: with ordered context, limited permissions, technically enforced checks and visible costs. In Dr. Sven Seiler's model it consists of seven layers – context, workflow, agent runtime, quality and gates, operations and learning loop, cost, and people and roles. The term “software factory” dates back to the 1970s; since 2026 BCG Platinion and Red Hat, among others, use it as “agentic software factory”.

How autonomously may an AI agent work in software development?

As far as verification carries. The agentic engineering model that Dr. Sven Seiler developed at Storm Reply with colleagues rates autonomy and verification separately per phase on five stages (assisted, delegated, supervised, autonomous, self-directing) and takes the minimum. Not every phase should reach stage 5: for design & architecture and review & acceptance the model recommends stage 3, because the human stays in the decision path.

What is agentic coding?

Agentic coding is programming with AI agents that edit files, run commands and use tools on their own instead of only making suggestions. Examples are Claude Code, GitHub Copilot in agent mode, Cursor and Kiro. Agentic engineering is the method teams use to apply this technique in a controlled, verifiable way.

What distinguishes Agentic Engineering from vibe coding?

Vibe coding is prompt → output → hope: no structure, no verification, no specification. Works for prototypes, breaks in production. Agentic Engineering inverts this: specifications as runtime component, six guardrails (CI / types / linting / architecture tests / behaviour tests / code quality with MCP), and clear responsibilities between human and agent.

Do I need Agentic Engineering if my team is small?

Especially then. In my experience, small teams with well-instrumented agents can keep pace with much larger teams — but only if the guardrails are in place. Without specs and tests the AI doesn’t scale, it produces tech debt at weekly cadence.

Which tools belong in an Agentic Engineering toolchain?

Coding agents (Claude Code, GitHub Copilot, Cursor, Kiro, Amazon Q Developer), MCP servers for external systems (DMS, ITSM, observability), static type systems (TypeScript, Rust, C#), architectural tests (ArchUnit and friends), code quality tools with MCP integration (SonarQube, CodeScene), and a strict trunk-based CI workflow.

What is the Model Context Protocol (MCP)?

MCP is an open standard for bidirectional, controlled connections between AI applications and external systems. Often described as “USB-C for AI.” In the Agentic Engineering context, MCP enables uniform feedback of quality reports, architectural constraints, and tool capabilities to agents — the foundation for autonomous refactoring.

Where does the term &ldquo;Agentic Engineering&rdquo; come from?

From Andrej Karpathy. He [proposed the term on X](https://x.com/karpathy/status/2019137879310836075) on 4 February 2026 as the successor to “vibe coding,” which he had coined a year earlier. Dr. Sven Seiler and Henning Teek elaborated it for enterprise practice in their April 2026 talk at the Agentic Shift Meetup in Dortmund and in the companion essay [“Coding Is Cheap, Software Is Not”](/articles/coding-is-cheap-software-is-not/en.html).

How do I measure success in Agentic Engineering projects?

Three metric families: **productivity** (release frequency, lead time, throughput), **quality** (defect rate, MTTR, test stability), and **risk** (coverage of critical paths, audit-trail completeness, compliance findings). Successful teams run these in parallel and conduct regular AI assessments — according to a [Gartner survey (November 2025)](https://www.gartner.com/en/newsroom/press-releases/2025-11-04-gartner-survey-finds-regular-ai-system-assessments-triple-the-likelihood-of-high-genai-value), organisations that regularly assess their AI systems are over three times more likely to report high GenAI value.

## Sources

[^1]: Gartner: “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027”, press release, 25 June 2025. [gartner.com](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027)

[^2]: Andrej Karpathy: post on X, 4 February 2026. [x.com](https://x.com/karpathy/status/2019137879310836075)

[^3]: JetBrains Research: “The State of Developer Ecosystem 2025”, October 2025. Survey of 24,534 developers, April to June 2025. [blog.jetbrains.com](https://blog.jetbrains.com/research/2025/10/state-of-developer-ecosystem-2025/)

[^4]: Cursor: “Past, Present, and Future” (Series D announcement, $2.3B at a $29.3B post-money valuation), 13 November 2025. [cursor.com](https://cursor.com/blog/series-d)

[^5]: Anthropic: “Anthropic acquires Bun as Claude Code reaches $1B milestone”, 3 December 2025. [anthropic.com](https://www.anthropic.com/news/anthropic-acquires-bun-as-claude-code-reaches-usd1b-milestone)

[^6]: Google Cloud / DORA: “2025 DORA Report: State of AI-Assisted Software Development”, study, 23 September 2025. Nearly 5,000 technology professionals worldwide, plus over 100 hours of qualitative data. [cloud.google.com](https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report)

[^7]: DX (Abi Noda, Laura Tacho): “Measuring AI code assistants and agents”, research paper on the DX AI Measurement Framework (dimensions: utilization, impact, cost). Undated. [getdx.com](https://getdx.com/research/measuring-ai-code-assistants-and-agents/)

[^8]: Augment Code: “Agentic Engineering Maturity Model: 5-Level Self-Assessment”, guide by Ani Galstian, 12 June 2026 (updated 18 June 2026). Five levels: Ad-Hoc, Standardized, Orchestrated, Systematic, Autonomous. [augmentcode.com](https://www.augmentcode.com/guides/agentic-engineering-maturity-model)

[^9]: Microsoft Learn: “Introduction to the Agentic AI adoption maturity model”, guidance, 31 March 2026. Five maturity levels (Level 100–500) and five capability pillars. [learn.microsoft.com](https://learn.microsoft.com/en-us/agents/adoption-maturity-model/)

[^10]: BCG Platinion: “The Agentic Software Factory”, insights article, 26 March 2026. Authors include Joachim Engesser, Axel Griewel, Sebastian Ley. [bcgplatinion.com](https://www.bcgplatinion.com/insights/the-agentic-software-factory)

[^11]: Red Hat Developer: “Trusted software factory: Building trust in the agentic AI era”, article by Meg Foley, 13 May 2026. [developers.redhat.com](https://developers.redhat.com/articles/2026/05/13/trusted-software-factory-building-trust-agentic-ai-era)

[^12]: Bain & Company: “From Pilots to Payoff: Generative AI in Software Development”, Technology Report 2025, 23 September 2025. [bain.com](https://www.bain.com/insights/from-pilots-to-payoff-generative-ai-in-software-development-technology-report-2025/)

[^13]: METR: “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, randomized controlled trial, 10 July 2025 (arXiv:2507.09089). 16 experienced open-source developers, 246 tasks. [metr.org](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)

[^14]: Veracode: “2025 GenAI Code Security Report”, study, 30 July 2025. Over 100 LLMs, code in Java, Python, C# and JavaScript. [veracode.com](https://www.veracode.com/blog/genai-code-security-report/)

[^15]: LeadDev: “The AI Impact Report 2025”, p. 13. Survey of 883 members of the LeadDev community, 30 May to 21 June 2025. [leaddev.com](https://leaddev.com/wp-content/uploads/2025/08/THE-AI-IMPACT-REPORT-2025-download__LDMO__.pdf)

## Deeper dives

### Coding Is Cheap, Software Is Not

The full essay on Agentic Engineering — era map, five-rung ladder, spec-as-runtime, six guardrails.

### AI-Powered Software Development 2025–2030

How generative AI and Agentic AI are transforming software development — adoption, governance, team roles, ROI.

### Talk: Coding Is Cheap (Agentic Shift, April 2026)

Talk with Henning Teek at the Agentic Shift Meetup in Dortmund, April 2026 (event page).
