How the layers link
Arrows show what calls or feeds what. Solid lines are calls; dashed lines return state or results. Hover a layer to trace its links; click it to jump to its definition below.
Hover a layer
Layers
| Layer | What it gives the template | Group | Across environments |
|---|
All links
| From | To | Relationship |
|---|
Simulation: an external model request, end to end
A simulated walkthrough of one request through every layer: plan, correct, build, compare, approve, score, and revert. The client, names and numbers are made up. Step through it, or press play.
acme_lookalike| Version | Stage |
|---|
| Job | Type | Env | Status |
|---|
| Step | Who | Action | Env | Result |
|---|
Environments for a small team
Dev, staging and prod are needed, but not every layer needs three copies. Each layer is one of four kinds, and each kind handles environments differently. Environments are about where platform code runs; they are not model stages.
Which stores are shared
| Store | Scope | Why |
|---|---|---|
| Model store | Shared | A model version is one object everywhere; stages are aliases on it. |
| Metadata store | Shared | Travels with the model version, so every environment's read path sees the same results. |
| Feature catalog | Shared | Feature metadata is attached to model versions. |
| Tool registry | Shared | Tool definitions are shared; access rules are set per environment. |
| Action log | Shared | One record of agent and operator actions; each entry names its environment. |
| Eval store | Shared | Eval runs and feedback attach to an agent version, which is the same object in every environment; each feedback entry names its environment. |
| Job store | Per environment | Runs belong to the environment that executed them. |
Rules
- Two cheap environments, one real. Dev can be mostly local; staging is prod at a smaller scale; only prod needs full capacity.
- Environment comes from config, never from code. One setting selects every per-environment resource.
- Where models are trained. Model training registers directly into the shared registry, from whichever environment ran it, so a model is uploaded once and never copied between registries. Dev is for changing the platform itself, not for holding models.
- Data rule. Dev and staging read prod input data (or a sample) and never write to prod data or prod outputs. Registering models and eval evidence is the one write they share, by design.
- Agent evals run before prod. The pre-release harness runs in dev or staging against fixture data; in-chat feedback is collected in every environment.
- One LLM interface across environments. Each tool call targets an explicit environment, and access controls and trigger levels are set per environment, so the same agent can freely read dev results and still ask before starting a prod job.
Model lifecycle
Promotion lives in the model registry. The registry owns one thing: which version of each model holds each stage. Promotion is the only way a model changes stage, any retained version can be promoted, and reverting is promoting a previous version. A human is always in the loop.
Stages
Who does what
| Part of promotion | Where it lives |
|---|---|
| State — which version holds each stage | Model registry (platform services) |
| Policy — the gates a model must pass | Each product line, versioned with its models |
| Evidence — eval runs against those gates | Run by orchestration; stored on the model version |
| Gate check — evidence vs. policy | The registry, at stage-change time; it refuses a change without passing evidence on that exact version |
| Decision — the approval | An analyst, after reviewing gains, lift, deciles, holdout performance and the audit |
| Consumption — scoring uses "production" | Orchestration resolves the alias through the registry; the read path reports on it |
Approval path
- Promotion moves an alias. Nothing is retrained or re-uploaded.
- Re-promoting a previous version still takes an analyst's sign-off, but its evidence already exists, so it can be approved immediately in an incident.
- Promotion doesn't change outputs already delivered; re-scoring affected batches is a separate orchestration step, found through the job store's model links.
Stores and lineage
Lineage is carried as metadata across the stores. Read in order, they form a chain from input data to delivered output.
| Store | What it tracks | Main readers |
|---|---|---|
| Model store registry record + artifact | Identity, version, stage and stage history. Test metrics and holdout evaluation. Training data references. The full pipeline config. The audit, including the analyst's approval (who, when, why). Pointer to the artifact, which holds the complete pipeline result. | Promotion, orchestration, scoring |
| Metadata store key:value components | The model summary stored as key:value components, each under a stable dotted key path (e.g. model_summary.roc_curve), plus an index listing every key with its type and size. Components include feature importances, coefficients, thresholds, ROC and gains curves, rank concentration and stability, coverage, business metrics, decile and ranking reports. The same key means the same thing in every version. | Read path, LLM interface, reporting, monitoring |
| Feature catalog | Feature metadata, not feature values: each feature's name, type, importance and tags, plus a feature group tying the set to its model. | Product lines, monitoring |
| Job store | Each run, in platform terms: workflow, input source, input-check results, expected count, model version, completeness check, delivery, and the engine's run ID. Status (pending through completed or failed) is copied from the workflow engine's state-change events; the engine owns run state, and the job store records it. | Operator UI, LLM, alerts |
| Tool registry | Tool definitions with version, parameter and return schemas, and execution strategy (function, module or API); plus access rules and trigger levels per environment. | LLM interface, agent framework |
| Action log | Every agent and operator action: who asked, on whose behalf, which tool, with what parameters, against which environment, the result, and links to any job or model it touched. | Analysts, audits, the LLM (its own history) |
| Eval store | Whether the agent did it well. Golden-set versions; eval runs per agent version with per-case scores, traces and the release verdict; and in-chat feedback (rating, flag, correction) linked to the action-log entries behind each answer. | Agent release gate, eval reviewers, monitoring |
LLM interface
One LLM interface serves every environment. It reads experiment results from the metadata store, starts jobs and checks their status, all through tools in the tool registry. The approach is guardrails, not handcuffs: the LLM gets real access, and the guardrails shape how it uses it.
Capabilities and guardrails
| Capability | What the LLM does | Guardrail |
|---|---|---|
| Read results | Lists experiments and versions, loads summary components, compares and plots them | Runs freely; every answer cites model ID, version and environment |
| Check job status | Reads run state from the job store, and step-level status, failure reasons and logs from the workflow engine | Runs freely |
| Start jobs | Launches builds and scoring runs through the tool registry, which hands them to orchestration | Same workflows and checks as the operator UI; explicit target environment; trigger level per environment (e.g. confirm before launching in prod) |
| Propose promotion | Assembles gains, lift, deciles, holdout results and the audit, and requests a stage change | Only an analyst's signed-in approval changes the stage |
| Every action is written to the action log. Access rules and trigger levels live in the tool registry and are enforced when a tool is executed. Every answer carries feedback controls, and every agent version passes the eval harness before release (see Agent evaluation). | ||
Reading results from experiments
execute_tool; the registry checks access and routes to the read path.| Read-path action | Used for |
|---|---|
get_details | Discover the action catalog instead of guessing |
list_models / search_models | Find experiments by name (fuzzy) |
list_model_versions | Enumerate the runs of one experiment |
get_model_summary | Compact first look: version, sizes, top features |
list_available_components | Discover which keys a version has |
load_specific_components | Fetch only what's needed: ROC, gains, deciles, coefficients |
load_model_metadata | Everything; last resort |
- Each tool does one job. Reporting tools read; job tools start and check runs; stage changes go through promotion.
- Keyed component reads. The agent reads the index to discover what exists, then fetches only the keys it needs, keeping context small and answers fast.
- Skills hold the contract. Each skill maps common requests to actions and documents the call and response shape.
- The registry holds the schema. Tool definitions carry parameter and return schemas and a version, so skills and tools can be checked against each other.
- Two status levels. The registry reports whether routing worked; the tool reports whether the action worked.
- Compare across versions. Because a key means the same thing in every version, a comparison is: list versions, read the same key from each, tabulate or plot.
- Monitoring as a tool. Drift results reach the agent through the tool registry, the same way metadata does.
- Taxonomy on demand. The agent reads the taxonomy as markdown, searches it through a hierarchical vector index, or queries it as an MCP server, so field meanings are looked up rather than guessed.
Agent evaluation
The action log records what the agent did. Evaluation records whether it did it well. There are two loops, and they share one golden set and one eval store. A pre-release harness gates every change to the agent, and feedback controls in chat catch what the harness missed.
What gets evaluated
| Area | What's measured | How it's scored |
|---|---|---|
| Retrieval | Taxonomy lookups and skill selection return the right fields and skills. Measured as recall@k and the rank of the first correct result on labeled queries, including whether near-target fields are flagged. | Deterministic, against labeled relevant sets |
| Tool use and trajectory | The right tool, parameters and environment at each step, and the right sequence overall: no skipped verification steps and no redundant calls. Confirmation is asked wherever the trigger level requires it. A fluent answer that skipped a check counts as a worse failure than a visible error. | Deterministic, against expected call traces; required steps must appear, in order |
| Grounding | No hallucinated numbers, fields or versions: every number in an answer matches a component the agent read, and every answer cites model ID, version and environment. | Numbers checked deterministically against tool results; an LLM judge checks for unsupported claims |
| Response quality | The answer is clear, complete and useful for the question asked. It states its assumptions, and it asks rather than guesses when a request is ambiguous. | LLM judge against a rubric, calibrated on analyst ratings |
| Task completion | Playbook tasks reach the expected end state: plan drafted, gates respected, the right job started. | State assertions in a dev sandbox, plus a rubric |
| Guardrails | The agent ignores instructions planted in content it reads. It never approves its own promotion, and it never launches in prod without confirmation. | Every case must pass; one failure blocks release |
| Cost and latency | Tokens, tool calls and time per task. | Compared with the current release; regressions are flagged but don't block by default |
Before release: an agent version is promoted like a model
Responsibilities split the same way as in model promotion: policy (the thresholds) is versioned in code, evidence (the eval run) is stored on the version, and a person decides.
In use: feedback in chat
v2026-09-30 against 2.7 for production v2026-08-01 (acme_lookalike, shared registry).
| Signal | Captured as |
|---|---|
| Helpful / not helpful | One click on any answer |
| Flag | A reason: wrong number, wrong field, wrong tool, should have asked, unsafe action |
| Correct | The right value or field, in the analyst's words |
| Implicit | Plan rejected, confirmation cancelled, question asked again |
| Each signal is stored in the eval store with the trace ID, agent version, skill and environment. | |
Closing the loop
Rules
- Golden sets are code. They are versioned, reviewed, and owned per skill or playbook, and every case is labeled with its area.
- Deterministic checks first, judges second. Anything that can be checked exactly, such as numbers, tool calls or retrieved fields, is checked exactly. LLM judges handle only what can't be.
- Judges are checked against people. Each LLM judge is calibrated on human-labeled cases, and re-checked whenever its model or rubric changes.
- Scores belong to an exact version. Changing any part of the agent version makes its scores stale until the suite runs again.
- Eval data follows the LLM data rules. Fixtures and captured traces carry no raw personal data.
Orchestration patterns
Orchestration receives job requests from the tool registry, the operator UI or arriving data, and runs them on a managed workflow engine and managed compute. The platform writes the workflow definitions and adds two checks specific to its data: input checks at the start and a completeness check at the end. Everything else (scheduling, provisioning, retries, cleanup, run state and events) belongs to the managed services.
Engine and compute options
| Option | Role | Good fit when |
|---|---|---|
| Airflow Amazon MWAA | Workflow engine: DAGs, schedules, data-aware triggers, retries, run history | The team knows Airflow, or workflows span systems beyond AWS |
| AWS Step Functions | Workflow engine: state machines with native Batch, EMR and SageMaker integrations | AWS-native, event-driven workflows with little to operate |
| AWS Batch | Compute: job queues, array jobs, retries, timeouts, state-change events | Container-based training and scoring tasks |
| Amazon EMR or EMR Serverless | Compute: managed Spark | Large, Spark-shaped data processing |
| Any engine plus any compute option gives the same system. Choose by the team's skills and the shape of the workloads. | ||
Built as code
The whole system is version controlled, not just the models. Infrastructure is defined in code (for example with the AWS CDK), so every service, store, queue, role and trigger has a reviewed, versioned definition, and an environment is that definition deployed with its own config.
What lives in version control
| Artifact | Defines | Released as |
|---|---|---|
| Infrastructure stacks | Platform services, orchestration, read path, monitoring, LLM interface hosting, roles and triggers | Deployed per environment, or once for shared resources |
| Shared libraries | Agent framework, ML framework, platform client, LLM knowledge | Tagged package versions |
| Workflow definitions | DAGs or state machines, their triggers, retry and timeout policies, and compute settings | Deployed per environment with orchestration |
| Tool definitions and access rules | Tool schemas, execution strategy, and trigger levels per environment | Deployed to the tool registry |
| Product-line code and gates | Model configurations and promotion gates | Versioned with the product line |
| Agent playbooks and skills | Staged workflows, approval gates, skill contracts | Reviewed and merged; one copy |
| Golden sets and eval thresholds | Golden tasks, labeled retrieval queries, judge rubrics, and release thresholds for the agent | Versioned with the playbooks and skills they test; run in CI on every change to the agent |
Models are the one thing deliberately not deployed from version control: they are data, versioned in the registry and moved by promotion. Their configs, gates and evidence are still traceable to versioned code.
Standards alignment
There is no single MLOps standard, and no formal standard yet for LLMs inside MLOps. There are widely used maturity models, governance frameworks, an emerging agent-standards effort, and de facto tooling. This tab maps the architecture to each.
Maturity models and governance frameworks
| Framework | How this architecture aligns |
|---|---|
| Google MLOps levels 0 manual · 1 pipeline automation · 2 CI/CD automation | Level 1–2. Pipelines are configuration on the ML framework and run by orchestration; packages and infrastructure are released through CI/CD (see Built as code). |
| Microsoft MLOps maturity levels 0–4 | Level 3: automated training, registry, lineage and deployment. Level 4's fully automatic promotion is deliberately not adopted: an analyst approves every promotion. |
| AWS MLOps phases initial · repeatable · reliable · scalable | Repeatable and reliable: infrastructure as code, separate environments from one definition, a shared registry, gated promotion and managed orchestration. |
| NIST AI RMF Govern · Map · Measure · Manage | Govern: access rules, trigger levels and approval policy, all in version control. Map: product lines with their own gates and audiences. Measure: input checks, eval evidence, drift monitoring, and agent evaluation before release and in use. Manage: promotion, re-promoting a previous version, and re-scoring. |
| ISO/IEC 42001 AI management system | Defined roles (the analyst approves), retained records (approvals on the model version, the action log), and a review loop (monitoring informs retraining decisions). |
| NIST AI Agent Standards Initiative launched February 2026 | Its first focus, agent identity, authorization and auditing, maps to three parts of the design: approvals made under the analyst's own identity, access rules enforced in the tool registry, and the action log. Treating read content as data, never instructions, addresses its prompt-injection concern. |
| Vibe coding to agentic engineering Google, The New SDLC With Vibe Coding, 2026 | The agentic-engineering end of the spectrum. Tests cover the deterministic parts (input checks, gates); evals cover trajectory, tool use and response quality. The harness (skills, tools, hooks, routing, context) is versioned, reviewed and evaluated in CI, and people own the architecture and the approvals. |
| Human oversight in-the-loop · on-the-loop | Human-in-the-loop for promotion: nothing changes stage without an analyst. Human-on-the-loop for runs: confirmation before prod launches, alerts and monitoring. |
Tooling and protocol interoperability
| Layer or store | Maps to |
|---|---|
| Model registry | MLflow Model Registry (versions, aliases such as "production", tags) or SageMaker Model Registry (model package versions with approval status such as pending manual approval / approved). The analyst approval maps directly to SageMaker's manual approval. |
| Metadata store | MLflow run metrics, parameters and artifacts; SageMaker model metadata and model cards. Keyed components map to logged artifacts addressed by path. |
| Orchestration | Workflow engine: Airflow (Amazon MWAA), AWS Step Functions or SageMaker Pipelines. Compute: AWS Batch, Amazon EMR (or EMR Serverless) for Spark, or SageMaker training and processing jobs. State-change events through EventBridge or Airflow callbacks. |
| Job store | MLflow runs or SageMaker Pipelines executions can hold the same record; either way, it carries the engine's run ID. |
| Tool registry | Model Context Protocol. Tool definitions are published in MCP schema format, and the registry exposes an MCP-compatible JSON-RPC 2.0 endpoint for listing and calling tools, so any MCP client can use the platform's tools. |
| Read path | The same role as MLflow's MCP server, which exposes the registry and experiment history to agents; either can sit behind the tool registry. |
| Agent framework and LLM knowledge | MCP client support for external tool servers; skills as markdown with front matter; the taxonomy served as markdown, a vector index or an MCP server. |
| Agent evaluation | MLflow GenAI evaluation and tracing, or Amazon Bedrock evaluations (including RAG evaluation); Ragas-style metrics for retrieval and grounding; OpenTelemetry GenAI semantic conventions for traces, so feedback and scores join on the same trace ID. |
| Built as code | AWS CDK (CloudFormation underneath); the same principles apply with Terraform or Pulumi. |
Where the design goes further
References
- Microsoft: MLOps maturity model
- An empirical guide to MLOps adoption (Google, Microsoft and AWS models compared)
- NIST AI RMF and ISO/IEC 42001
- NIST launches AI Agent Standards Initiative
- CSA: Agentic AI governance and NIST standards
- AWS: AI agents with SageMaker AI and MCP
- AWS: Agents with SageMaker AI models and MLflow
- Osmani, Saboo and Kartakis: The New SDLC With Vibe Coding (Google, May 2026)
Decisions to make
Choices the design leaves to the implementer. Each has a reasonable default.
| Decision | Options and default |
|---|---|
| Gates per product line | Which metrics and thresholds a candidate must pass before the analyst sees it. Default: start from the metrics analysts already review (gains, lift, deciles, holdout gap) and the audit flags. |
| LLM access controls and trigger levels | Which actions run directly, which ask for confirmation and which need approval, per environment. Default: reads and status checks run directly; starting jobs asks for confirmation in prod; promotion needs analyst approval. |
| LLM data and safety rules | What data may be sent to the LLM provider, and how content read from files and metadata is treated. Default: no raw personal data in prompts; content the agent reads is treated as data, never as instructions. |
| Agent release thresholds | What an agent version must score before release. Default: guardrail cases must all pass; retrieval, trajectory, grounding and response-quality scores may not fall below the current release by more than a set margin; cost and latency regressions are flagged but don't block. |
| Golden set ownership | Who writes and maintains the eval cases. Default: each skill or playbook owner owns its cases; a starting set is seeded from the analysts' most common requests and from known failures. |
| Feedback triage | Who reviews flags and corrections, and how quickly. Default: weekly review by the skill owner; guardrail flags are reviewed the same day. |
| LLM judges | Where a judge is used, and which model it runs on. Default: only for checks that can't be done exactly; a different model from the agent's; calibrated on human labels before use. |
| Input checks | What is checked before training and scoring. Default: schema against the data dictionary, row counts against expectations, and null rates and distributions against a reference. |
| Retraining trigger | What prompts a new candidate. Default: an analyst decides from the drift dashboard; the agent can propose. |
| Outcome feedback | How real outcomes, when they arrive, flow back. Default: recorded against the model version in the metadata store and shown in monitoring. |
| Output delivery | Who owns getting outputs to consumers. Default: orchestration delivers and records the delivery on the job. |
| Deploy automation | Whether infrastructure deploys run by hand from reviewed code or automatically on merge. Default: automatic to dev and staging; prod deploys after review of the infrastructure diff. |
| Retention and backup | How long versions are kept, and how the shared registry is protected. Default: keep every version that was ever in production; back up the registry and its artifacts. |
| Cost limits | Caps on compute and LLM use. Default: a maximum capacity per compute environment or queue, a per-run concurrency limit, and a per-session LLM budget, with alerts. |
| Model routing | Which LLM model each step uses. Default: the most capable model for planning and analysis; smaller, cheaper models for routine steps such as taxonomy lookups, summarizing tool output and simple status questions; judges on a different model from the agent's. Routing rules are part of the agent version, so changing them re-runs the eval. |
| Workflow engine and compute | Which managed services run orchestration. Any of them gives the same system: Airflow (MWAA) or Step Functions as the engine, and AWS Batch, EMR or SageMaker jobs as compute. Default: choose by team skills and workload shape, such as Airflow where the team already runs it, Step Functions for AWS-native event-driven work, Batch for container tasks, and EMR for Spark. Whatever is chosen, keep the input checks, the completeness check and the job store, so the LLM tools stay the same. |
| Online serving | Batch only, or an optional real-time serving layer. Default: batch only until a product line needs real-time scoring. |