Skip to main content

Command Palette

Search for a command to run...

The Agent-Ready Enterprise

How AAIF, AGENTS.md & agentgateway are changing software engineering

Updated
•22 min read•View as Markdown
The Agent-Ready Enterprise
D
I'm Ayyanar Jeyakrishnan ; aka AJ. With over 21 years in IT, I'm a passionate Multi-Cloud Architect specialising in crafting scalable and efficient cloud solutions. I've successfully designed and implemented multi-cloud architectures for diverse organisations, harnessing AWS, Azure, and GCP. My track record includes delivering Machine Learning and Data Platform projects with a focus on high availability, security, and scalability. I'm a proponent of DevOps and MLOps methodologies, accelerating development and deployment. I actively engage with the tech community, sharing knowledge in sessions, conferences, and mentoring programs. Constantly learning and pursuing certifications, I provide cutting-edge solutions to drive success in the evolving cloud and AI/ML landscape.

The future of enterprise software engineering is not one agent, one model, or one coding assistant. It is an agent-ready enterprise where knowledge, context, execution, control and evaluation are independently governed.

ABSTRACT — The short version

Software engineering is moving from AI-assisted development to agentic software engineering. Coding agents now navigate repositories, change many files, call tools, run tests and complete multi-step tasks with growing autonomy.

But enterprise software is overwhelmingly brownfield: applications spread across many repositories, with legacy interfaces, undocumented business rules, proprietary frameworks, security constraints, regulatory obligations and established pipelines. In that environment, asking every agent session to rediscover the application is expensive and inconsistent. The opposite move — pouring the whole knowledge base into AGENTS.md — fails too: research on repository context files found they can raise inference cost and, when padded with unnecessary requirements, fail to improve or even reduce task success.

This article proposes an Agent-Ready Enterprise architecture built on a six-stage lifecycle that produces evidence, not just code:

DISCOVER → KNOW → COMPILE CONTEXT → EXECUTE → CONTROL → EVALUATE ⇒ EVIDENCE

Each application maintains two auxiliary repositories: a Knowledge Hub (what the application means) and a Test & Evaluation Hub (what the application must prove). A context compiler turns hub knowledge into compact, repository-level AGENTS.md. Coding, discovery and evaluation may use different models, and agentgateway enforces which of them may see which code. Changes merge only when evidence satisfies the contract.

Ten ideas to take away

  1. Ask “Is our enterprise ready for an agent?”, not “Which assistant should we buy?”

  2. In brownfield, the first AI task is discovery, not implementation.

  3. The application, not the repository, is the unit of knowledge.

  4. Knowledge carries a confidence state and provenance; unverified facts are flagged, not asserted.

  5. AGENTS.md is a local operating contract — compiled, not dumped.

  6. Discovery, coding and evaluation are different jobs and can use different models.

  7. The agent picks the task; policy picks the permitted intelligence — including for the evaluator.

  8. Every contract change updates both what the agent knows and how its work is judged.

  9. Enterprise Guardrails will be defined as skills or MCP[Eg for Language, Artifacts, Approved Libraries ] and referred Hard gates first, weighted scores second; the agent never grades itself.

  10. Prove the harness works with a counterfactual: the same tasks with and without it.

The Agent-Ready Enterprise

“Make the enterprise ready for agents.”

From a tangled brownfield estate to a governed agentic engineering system: discover, know, compile context, execute, control, evaluate.

Part I · The problem

1) From AI-assisted coding to agent-ready engineering

The first wave of AI coding was about acceleration. A developer described a change, an assistant produced code, and the developer reviewed it. The human held the context; the model filled in keystrokes.

The agentic model shifts who holds the context. An agent inspects repositories, decides which files to change, calls tools, runs tests and iterates across many steps. Recent work on Spec-Driven Development treats specifications as a contract between humans and agents, governed by a technical and methodological harness [2]. Once the agent is doing the exploring, the quality of what it can find — and what it is allowed to touch — becomes the constraint.

So the enterprise question changes. Instead of asking which coding assistant developers should use, ask:

Is our enterprise ready for an agent?

A production agent working on a real enterprise system needs to know:

  • what the application does and which repositories make it up;

  • which modules a change affects, and which business rules must not change;

  • which APIs, events and data stores are involved;

  • which security and regulatory constraints apply;

  • which models are permitted to process this code;

  • which tests define acceptable behaviour;

  • how the generated change will be independently verified.

None of that is a property of the model. It is a property of the enterprise. That is why the next discipline is not AI-assisted development but agent-ready engineering.

2) Brownfield is the real enterprise problem

Consider a large enterprise product with 20 applications, each spanning 5–12 repositories: services, front ends, batch jobs, data pipelines, APIs, stored procedures, infrastructure, configuration, integration adapters and legacy components. A single business capability routinely crosses several of them.

Product → Application → Business Capability → Modules → 

Repositories
        → APIs → Data → Business Rules → Tests → Deployment

A requirement like “change the pricing calculation for Product X” is therefore not a repository edit. The agent may need to understand the pricing engine, customer segmentation, the contract service, database rules, audit events and downstream consumers.

If every task begins by rediscovering this graph, the organisation pays the discovery cost again and again. Worse, two agents — or the same agent on two days — can form two different interpretations of the same application.

Context divergence — The code is shared. The agent’s understanding of it is not. Without a durable, shared account of what the application means, every agent session is a private, unreviewed architecture opinion.

Brownfield adds a second problem that greenfield hides: much of what matters was never written down. Business rules live in stored procedures, conventions live in reviewers’ heads, and the reasons behind constraints live in old incident tickets. An agent-ready architecture has to recover that knowledge, record how sure it is of it, and keep it current.

Part II · The architecture

1) The six-stage lifecycle

The architecture separates six concerns, each with its own question and its own artifact. Evidence is not a seventh stage; it is what the lifecycle produces.

# Stage Question it answers Primary artifact
1 Discover What does this brownfield application actually do? Application Contract
2 Know What does the enterprise know about it, and how sure is it? Knowledge Hub / OKF
3 Compile context What does this agent need to know, here, for this change? AGENTS.md
4 Execute How should the change be implemented? Coding agent + code patch
5 Control Which models, tools and data paths are permitted? agentgateway policy
6 Evaluate Did the change satisfy the contract — and can we prove it? Test & Evaluation Hub → evidence record

Six Stages of the Agent-Ready Lifecycle

An engineering lifecycle, not a product diagram.

Each stage maps to one artifact; a feedback loop returns from Evaluate to Know.

Stage 1: Discover: build the contract before writing code

For brownfield software, the first AI task should usually be discovery. Point a strong reasoning model at the source repositories, architecture, APIs, schemas, deployment definitions, configuration, dependencies, tests, documentation, business rules, security controls, ownership and defect history. The goal is not code. The goal is an application contract.

Application Contract
├── Purpose and business capabilities
├── Architecture and modules
├── Repository map (which repo implements which module)
├── API, event and data contracts
├── Business rules (with source locations)
├── Security and regulatory constraints
├── Deployment model
├── Ownership
├── Validation requirements
└── Open questions (what discovery could not determine)

Revised — Discovery output must include an explicit open questions list. What a model could not determine is as valuable as what it could — it is the queue for human owners to resolve, and it stops guesses from being promoted to facts.

Discovery is also expensive — a strong model reading hundreds of repositories is a real line item. Run it once per application, then incrementally on change, rather than per task. That is the economic argument for the next stage.

STAGE 2 - Know the Application Knowledge Hub

Every application gets a dedicated auxiliary repository — its durable, reviewable memory:

application-orders-knowledge-hub/
├── okf/                 # live knowledge graph for the application
│   └── application.okf
├── application-card/
│   └── application.yaml
├── modules/
│   ├── pricing.yaml
│   ├── order.yaml
│   ├── customer.yaml
│   └── fulfillment.yaml
├── contracts/           # api/, events/, data/
├── dependencies/
├── security/
├── ownership/           # CODEOWNERS-style, real people or teams
├── provenance/          # where each fact came from
└── compiler/            # context-compilation rules (Section 6)

The key idea: the application, not the repository, is the unit of knowledge. A pricing capability might span pricing-service, order-service, customer-service, pricing-database and event-processing. The hub represents that as one capability, with the cross-repository relationships made explicit in the live OKF — the application’s machine-readable graph of structure, modules, repositories, APIs, dependencies, business concepts, policies, ownership and provenance.

Knowledge has to earn trust

A hub populated entirely by a discovery model is a hypothesis, not a source of truth. Compiling it straight into agent instructions would give model guesses the authority of policy. So every fact carries a confidence state:

State Meaning How it is reached Use in context
inferred Proposed by discovery; plausible, unconfirmed Model output with source pointers Flagged as unverified, or excluded
verified Checked against code or a passing test Automated check or test in the Evaluation Hub Stated as fact
owner-attested Confirmed by the accountable owner Reviewed PR to the hub Stated as a constraint
stale Source changed since last verification Drift check in CI Excluded until re-verified

Watch for this — If most module cards sit at inferred for months, the hub is not a knowledge layer yet — it is an unreviewed model output with a nicer folder structure. Track the share of verified and attested facts as a first-class health metric.

Provenance is not optional

Every fact records where it came from: repository and commit, file and line range, test identifier, or reviewer and date. Provenance is what lets a drift check mark a fact stale when its source changes, and what lets an auditor later ask why an agent believed something.

STAGE 3 · Compile context: don’t turn AGENTS.md into the knowledge base

This is where discipline matters most. AAIF’s guidance describes AGENTS.md as a lightweight Markdown format and recommends including what actually changes agent behaviour: project structure, commands, conventions, testing instructions and security considerations [3]. ETH Zurich’s evaluation of AGENTS.md-style context files found they did not generally improve task success in their experiments while raising inference cost by more than 20%, and the authors caution specifically against unnecessary requirements [1].

More context is not automatically better context.

The context compiler

I call the transformation from enterprise knowledge to agent-ready instructions context compilation. Thoughtworks describes context engineering for coding agents as curating what the model sees to improve its results [4]; the compiler makes that curation a repeatable, reviewable capability instead of a craft each team reinvents.

The compiler is driven by rules that live in the hub and are reviewed like code:

compiler/pricing-engine.yaml

compile:
  target: pricing-engine/AGENTS.md
  include:
    module_cards: [pricing]
    dependencies: direct          # not transitive
    constraints: severity >= must
    commands: from_build_manifest # never hand-copied
    security: [classification, data_handling]
  exclude:
    - internals of other modules  # link, don't inline
    - design history and rationale
  confidence:
    state_as_fact: [verified, owner-attested]
    inferred: flag_as_unverified
    stale: exclude
  budget:
    max_lines: 150
  provenance:
    stamp_hub_commit: true        # AGENTS.md records which hub version produced it
  recompile_on:
    - hub change touching pricing
    - build manifest change

And the output stays small:

pricing-engine/AGENTS.md

# Repository context
<!-- compiled from application-orders-knowledge-hub@7c03be — do not edit by hand -->

## Purpose
Pricing calculation service for enterprise orders.
Application: application-orders · Module: Pricing

## Important dependencies
- Customer eligibility API
- Contract rules service
- Pricing database

## Commands
- Build: ./gradlew build
- Unit tests: ./gradlew test
- Integration tests: ./gradlew integrationTest

## Constraints
- Do not bypass PricingPolicy.
- Every pricing change emits an audit event.
- Do not change public API schemas without updating contracts/api.
- (unverified) Discount rounding may be applied in contract-service; confirm before relying on it.

## Security
- Customer pricing data is confidential.
- Restricted module: only internal models are permitted (enforced by agentgateway).

## Definition of done
- Unit tests updated; contract and regression suites pass.
- Evaluation threshold met; evidence record complete.

Design rule — A hand-edited AGENTS.md in a compiled repository is drift by definition. Stamp the hub commit into the file and fail CI when the stamp and the current hub disagree for the relevant module.

Source knowledge drawn as rich and comprehensive; AGENTS.md drawn as intentionally compact.

STAGE 4 · Coding with a different model

With discovery done and context compiled, the coding agent can work — and it does not need to use the discovery model. The implementation model might be an AWS Bedrock model, an approved external model, an internal model, a code-specialised model or a local open-weight model.

Discovery model → Application Contract → Knowledge Hub → Context Compiler
  → AGENTS.md → Coding agent (implementation model) → Code patch

The application should not be architecturally coupled to a provider. Models change quickly, costs differ by stage, regional availability varies, data classification restricts choices, and some workloads must stay on-premises. The model layer should be replaceable and policy-controlled.

Where Kiro and other developer experiences fit

Kiro, Cursor or any other coding agent can own the developer experience: specification, requirements, design, tasks, implementation and agent interaction. None of them should become the enterprise knowledge architecture.

Standardise knowledge, context, model policy, security, evaluation and evidence. Let developers choose — and change — the tool they drive it with.

Human + Agent

The engineer’s work moves from typing every line to defining intent, architecture, constraints and verification.

STAGE 5 · Agentgateway as the policy boundary

Without a boundary, a coding agent talks directly to Model A, Model B, Bedrock, an on-premises model and an external provider. That leaves an ungoverned question: who decides which model may process this repository?

Agentgateway is an AAIF-hosted open-source gateway for AI and agent traffic — LLM inference, MCP, A2A, HTTP and gRPC — with authentication, authorisation, routing, observability, governance and model independence [5]. That makes it a natural enterprise control point. AAIF has also shown it fronting self-hosted models with authentication, token budgets, routing, failover and usage metrics [6], which is exactly what an enterprise needs to keep sensitive workloads on internal models.

Model policy as code

Classify modules, not just applications, and let the gateway enforce the result:

policy/payments.yaml

application: payments
modules:
  customer-portal:
    classification: internal
    allowed_models: [bedrock-approved, internal]
  payment-processing:
    classification: confidential
    allowed_models: [bedrock-approved, internal]
  fraud-detection:
    classification: restricted
    allowed_models: [internal]
defaults:
  unknown_module: deny           # unclassified code is not sent anywhere
  token_budget_per_task: 2_000_000
  audit: full_request_metadata

The gateway decision uses identity, application, repository, module, data classification, model policy and budget, and writes an audit record either way:

fraud-detection  →  internal model   ALLOW
fraud-detection  →  external model   BLOCK  (classification: restricted)

Revised — Policy applies to every agent that reads the code — discovery, coding and evaluation alike. An “independent” evaluator that sends restricted code to an unapproved provider is a data-handling breach with a good intention. For restricted modules, independence must come from a different internal model or a heavier deterministic test weight, not from an external judge.

Model diversity becomes an engineering control

Stop asking only “which is the best model?” Ask “which model is appropriate for this job, this data and this risk level?” Discovery maximises understanding; coding maximises implementation efficiency; evaluation maximises independent verification; internal models satisfy security and compliance. A multi-model architecture is not a technology preference — it is a control, the same way segregation of duties is a control.

Model Specialisation

“Model diversity becomes an engineering control.”

Discovery, coding, evaluation and internal model lanes connected through a policy gateway.

STAGE 6 · Evaluate: the Test & Evaluation Hub

Every application gets a second auxiliary repository. It is not just another test repo; it is the application’s executable validation contract, answering one question: how do we prove a proposed change still conforms to what the application is supposed to do?

application-orders-test-evaluation-hub/
├── validation-specs/
├── contract-tests/
├── regression-tests/
├── behavioral-scenarios/
│   ├── visible/        # may be referenced in agent context
│   └── held-out/       # never compiled into context; used only at evaluation
├── security-tests/
├── evaluation-rubrics/
├── validation-agent/
├── evaluation-judge/
├── scoring/
└── merge-gates/

Revised — Held-out scenarios. If every test the evaluator uses is visible to the coding agent, the agent can optimise toward the tests rather than the behaviour. Keep a held-out set that is never compiled into context, and rotate it.

The contract-change propagation rule

Suppose the business says: enterprise customers now receive a different pricing calculation. The requirement cannot stop at a specification document.

Every meaningful contract change changes both what the agent knows and how the generated code is evaluated.

A requirement is not done because the specification changed. It is done when the knowledge representation and the validation contract have changed with it — and CI should enforce that a knowledge change to a business rule without a matching evaluation change is flagged for review.

Independent evaluation: never let the agent grade itself

Vercel’s guidance on evaluating coding agents treats the harness, not the model, as the unit of evaluation, and highlights context selection, model choice across steps and independent verification [7]. Microsoft’s evaluation guidance recommends evaluation-driven development: define success before building, write test cases early, set measurable goals, surface assumptions and maintain regression tests [8]. Together they give the Evaluation Hub a wider remit than unit testing: it is the application’s behavioural contract.

Validation agent and evaluation judge

Agent patch
  → Validation specification
  → Validation agent
       ├── select or generate tests
       ├── execute the candidate
       ├── run deterministic checks
       ├── collect evidence
       └── invoke the evaluation judge → score + explanation
  → Merge decision

The judge considers functional correctness, contract compliance, regression safety, API compatibility, security constraints, behavioural fidelity, completeness, quality and unintended side effects. Senior SWE-Bench describes a similar pattern, where a validation specification drives behavioural verification and a judge evaluates the evidence [9]. The enterprise adaptation: every application owns its validation contract.

Ranking candidates: hard gates first, scores second

Evaluation need not be binary, but a weighted average alone can hide a disqualifying failure. Apply non-negotiable gates first, per-dimension floors second, and only then rank.

merge-gates/pricing.yaml

merge_gate:
  hard_gates:                 # any failure = reject, no score can compensate
    contract_tests: pass
    security_tests: pass
    regression_suite: pass
    evidence_record: complete
  scoring:
    weights: {functional: 0.30, contract: 0.30, security: 0.20, regression: 0.20}
    per_dimension_floor: 80
    minimum_overall: 90
  human_review_required_when:
    - classification == restricted
    - public_api_changed
    - overall < 95
Candidate Functional Contract Security Regression Weighted Decision
A 96 98 99 97 97.4 Eligible · rank 1
B 94 88 96 91 92.0 Eligible · human review
C 97 71 89 84 85.0 Rejected · contract floor

Candidate C has the best functional score and would look competitive on a leaderboard, but it breaks the contract. That is precisely the change an unguarded agent pipeline would merge. Scores are illustrative; weights and floors are set per application.

The Application Test & Evaluation Hub

“Every meaningful contract change changes both what the agent knows and how the generated code is evaluated.”

Validation spec, contract, regression and behavioural tests, validation agent, judge, score and evidence, merge gate: PASS → merge allowed, FAIL → merge blocked.

Part III · Putting it together

Section 1: The complete brownfield workflow

Eg Enterprise pricing Application

An enterprise pricing application spans pricing-api, pricing-engine, customer-service, contract-service, pricing-database, audit-service and deployment-infra. The business changes one rule: enterprise customers with contract type X receive a new pricing calculation.

Stage What happens
Discover Already done for this application; an incremental pass confirms the rule touches pricing engine → contract service → customer segment → audit service.
Know The rule is added to the pricing module card and the OKF with provenance, as owner-attested once the product owner approves the hub PR.
Evaluate (prepared) The Evaluation Hub adds the new scenario, contract tests, regression and negative cases, and a held-out variant the agent will not see.
Compile context pricing-engine/AGENTS.md is recompiled; only the constraints that matter to this change appear.
Execute The coding agent implements the change with a permitted coding model.
Control agentgateway sees pricing-engine is restricted: external models blocked, internal model allowed — for the coding agent and the evaluator alike.
Evaluate (run) Deterministic suites run; a different internal model judges behavioural fidelity; an evidence record is produced.
Merge Contract, regression and security pass; score ≥ threshold; restricted classification triggers human review; approver signs.

The evidence record

For an important automated change, the organisation should be able to answer, from one record: what changed, why, from what context, by which model, under which policy, tested how, judged by whom, and approved by whom.

evidence/PR-4821.json

{
  "change_id": "PR-4821",
  "requirement": "PRC-219: contract type X pricing for enterprise customers",
  "application": "enterprise-pricing",
  "repositories": ["pricing-engine", "contract-service"],
  "context": {
    "agents_md": "pricing-engine/AGENTS.md",
    "knowledge_hub_commit": "7c03be",
    "facts_used_unverified": 0
  },
  "models": {
    "coding": "internal-code-model",
    "evaluation": "internal-eval-model (different family)"
  },
  "policy": {
    "gateway_policy": "payments/pricing-engine@v12",
    "classification": "restricted",
    "decision": "allow: internal only",
    "blocked_attempts": 0
  },
  "tests": {
    "contract": "pass (42/42)", "regression": "pass (318/318)",
    "security": "pass", "held_out_scenarios": "pass (6/6)"
  },
  "evaluation": { "weighted_score": 96.1, "threshold": 90, "judge_summary": "…" },
  "decision": "merge_allowed",
  "human_approver": "pricing-owners"
}

SECTION 2: The six contracts enterprises should standardise

Earlier drafts listed five contracts in one place and a different five in another. The unified set is six — one per concern, plus the evidence that ties them together:

Contract Artifact Standardises
Specification Requirement / spec Intent, acceptance criteria, scope
Knowledge Knowledge Hub / OKF Application identity, modules, dependencies, APIs, capabilities, ownership, security, provenance, confidence
Context AGENTS.md + compiler rules Repository-local operating instructions and how they are derived
Control agentgateway policy Which model may process which data; internal-only workloads; providers; budgets; routing; audit
Validation Test & Evaluation Hub Required tests, regression suites, held-out scenarios, rubrics, scoring, merge gates
Evidence Evidence record What changed, why, with what context, model, policy, tests, judge and approver

SECTION 3: The closed feedback loop

The architecture is not linear. Production teaches it.

  • Production defects become regression cases.

  • New architecture becomes application knowledge.

  • New security requirements become model policy.

  • New coding conventions become compiler rules and repository guidance.

  • New business rules become both knowledge and evaluation criteria.

  • Evaluator disagreements with humans become rubric fixes.

AAIF as the open foundation — and the enterprise agent harness

AAIF is better presented as a set of open building blocks than as a product catalogue. It describes itself as a neutral home for open standards and open-source projects that form infrastructure for agentic systems [10]. Relevant pieces include AGENTS.md for repository guidance, MCP for tool and context interoperability, A2A for agent-to-agent interoperability, agentgateway for traffic governance, and goose as an open agent runtime.

How should an agent operate here?

The context contract, compiled per repository.

What may the agent access, with which intelligence?

The control contract, enforced per call.

Closing

The biggest mistake an organisation can make is to treat AI coding as a tool rollout. “Which assistant should we standardise on?”, “Which model should everyone use?”, “How many developers can one agent replace?” — those questions start too low in the stack.

Better questions:

  • What does our application actually know — and how sure are we?

  • What context does an agent really need for this change?

  • What should the agent be allowed to access, and which model may see this code?

  • How do we independently verify the generated change?

  • Can we prove why the change was allowed into production?

  • Can we show the harness makes outcomes better than not having it?

That is the difference between AI-assisted development and an agent-ready enterprise. Discover the brownfield system. Create a durable application contract. Keep knowledge in a living, provenance-backed hub. Compile only relevant context into AGENTS.md. Let coding agents use models suited to implementation. Use agentgateway to enforce model, tool and data policy for every agent. Maintain an independent Test & Evaluation Hub. Merge only when the evidence satisfies the contract.

Give the agent the right knowledge.

Give it only the context it needs.

Give it only the models and tools it is allowed to use.

And make it prove that its work satisfies the contract.

REFERENCES — Sources

  1. Thibaud Gloaguen et al. — Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? arxiv.org/abs/2602.11988

  2. Jessica Diaz et al. — Spec-Driven Development for Agentic Software Engineering: Harnessing Human-Agent Teamwork. arxiv.org/abs/2609.00252

  3. Agentic AI Foundation — Writing an Effective AGENTS.md. aaif.io/blog/writing-an-effective-agents-md

  4. Birgitta Böckeler / Thoughtworks — Context Engineering for Coding Agents. martinfowler.com

  5. Agentic AI Foundation — agentgateway Joins AAIF as an Open Gateway for Agentic AI Infrastructure. aaif.io

  6. Agentic AI Foundation — Putting a Front Door on Local AI with agentgateway. aaif.io

  7. Vercel — How to Evaluate AI Coding Agents. vercel.com

  8. Microsoft Learn — Agent Evaluation Overview / Evaluation-Driven Development. learn.microsoft.com

  9. Senior SWE-Bench — How Validation Works. senior-swe-bench.snorkel.ai

  10. Agentic AI Foundation — AAIF Blog and Project Ecosystem. aaif.io/blog

3 views