← WhoIsJaneTurner Research

MLOps reference architecture with an LLM interface

Generalized from a working Data Science platform · version 1.2 · October 4, 2026

How the layers link

Arrows show what calls or feeds what. Solid lines are calls; dashed lines return state or results. Hover a layer to trace its links; click it to jump to its definition below.

LLM interface Shared libraries Platform Product lines — call   - - - returns state or results

Hover a layer

Its incoming and outgoing links appear here.

Layers

LayerWhat it gives the templateGroupAcross environments
All links
FromToRelationship

Simulation: an external model request, end to end

A simulated walkthrough of one request through every layer: plan, correct, build, compare, approve, score, and revert. The client, names and numbers are made up. Step through it, or press play.

Model registry · acme_lookalike
VersionStage
Job store
JobTypeEnvStatus
Action log
StepWhoActionEnvResult

Environments for a small team

Dev, staging and prod are needed, but not every layer needs three copies. Each layer is one of four kinds, and each kind handles environments differently. Environments are about where platform code runs; they are not model stages.

Built once, versioned
Agent framework, ML framework, platform client, LLM knowledge, agent playbooks. Published or reviewed once; each environment pins the version it uses.
Deployed per environment
Orchestration (the managed workflow engine and its compute), model monitoring, and the job store. Same infrastructure definition, different config.
Shared across environments
LLM interface, model registry (model, metadata and feature stores), tool registry, action log, eval store. One instance; every call and record names its environment.
Product lines
Their code is versioned like any library. Their models are not deployed per environment; they move through registry stages.

Which stores are shared

StoreScopeWhy
Model storeSharedA model version is one object everywhere; stages are aliases on it.
Metadata storeSharedTravels with the model version, so every environment's read path sees the same results.
Feature catalogSharedFeature metadata is attached to model versions.
Tool registrySharedTool definitions are shared; access rules are set per environment.
Action logSharedOne record of agent and operator actions; each entry names its environment.
Eval storeSharedEval runs and feedback attach to an agent version, which is the same object in every environment; each feedback entry names its environment.
Job storePer environmentRuns belong to the environment that executed them.

Rules

  • Two cheap environments, one real. Dev can be mostly local; staging is prod at a smaller scale; only prod needs full capacity.
  • Environment comes from config, never from code. One setting selects every per-environment resource.
  • Where models are trained. Model training registers directly into the shared registry, from whichever environment ran it, so a model is uploaded once and never copied between registries. Dev is for changing the platform itself, not for holding models.
  • Data rule. Dev and staging read prod input data (or a sample) and never write to prod data or prod outputs. Registering models and eval evidence is the one write they share, by design.
  • Agent evals run before prod. The pre-release harness runs in dev or staging against fixture data; in-chat feedback is collected in every environment.
  • One LLM interface across environments. Each tool call targets an explicit environment, and access controls and trigger levels are set per environment, so the same agent can freely read dev results and still ask before starting a prod job.

Model lifecycle

Promotion lives in the model registry. The registry owns one thing: which version of each model holds each stage. Promotion is the only way a model changes stage, any retained version can be promoted, and reverting is promoting a previous version. A human is always in the loop.

Stages

Before training
Input checks
Schema, row counts and distributions checked against expectations. Failures stop the run and are recorded on the job.
Stage
Candidate
Registered, with eval evidence attached and gates evaluated. Waiting for the analyst.
Stage
Production
An analyst approved the metrics. Scoring resolves this alias.
Stage
Previous
Retained with its artifacts and evidence; can be promoted again.

Who does what

Part of promotionWhere it lives
State — which version holds each stageModel registry (platform services)
Policy — the gates a model must passEach product line, versioned with its models
Evidence — eval runs against those gatesRun by orchestration; stored on the model version
Gate check — evidence vs. policyThe registry, at stage-change time; it refuses a change without passing evidence on that exact version
Decision — the approvalAn analyst, after reviewing gains, lift, deciles, holdout performance and the audit
Consumption — scoring uses "production"Orchestration resolves the alias through the registry; the read path reports on it

Approval path

Approval the agent can't fake
The agent assembles the metrics and requests promotion. The approval itself is a signed-in action under the analyst's identity, sent from the operator UI or an approve control in the LLM interface, straight to the registry. It's never a message the agent relays. The registry records who approved, when and why.
Rules
  • Promotion moves an alias. Nothing is retrained or re-uploaded.
  • Re-promoting a previous version still takes an analyst's sign-off, but its evidence already exists, so it can be approved immediately in an incident.
  • Promotion doesn't change outputs already delivered; re-scoring affected batches is a separate orchestration step, found through the job store's model links.

Stores and lineage

Lineage is carried as metadata across the stores. Read in order, they form a chain from input data to delivered output.

Input data + config→Input checks→Model version→Features→Evaluation + summary→Approval→Runs→Outputs
StoreWhat it tracksMain readers
Model store
registry record + artifact
Identity, version, stage and stage history. Test metrics and holdout evaluation. Training data references. The full pipeline config. The audit, including the analyst's approval (who, when, why). Pointer to the artifact, which holds the complete pipeline result.Promotion, orchestration, scoring
Metadata store
key:value components
The model summary stored as key:value components, each under a stable dotted key path (e.g. model_summary.roc_curve), plus an index listing every key with its type and size. Components include feature importances, coefficients, thresholds, ROC and gains curves, rank concentration and stability, coverage, business metrics, decile and ranking reports. The same key means the same thing in every version.Read path, LLM interface, reporting, monitoring
Feature catalogFeature metadata, not feature values: each feature's name, type, importance and tags, plus a feature group tying the set to its model.Product lines, monitoring
Job storeEach run, in platform terms: workflow, input source, input-check results, expected count, model version, completeness check, delivery, and the engine's run ID. Status (pending through completed or failed) is copied from the workflow engine's state-change events; the engine owns run state, and the job store records it.Operator UI, LLM, alerts
Tool registryTool definitions with version, parameter and return schemas, and execution strategy (function, module or API); plus access rules and trigger levels per environment.LLM interface, agent framework
Action logEvery agent and operator action: who asked, on whose behalf, which tool, with what parameters, against which environment, the result, and links to any job or model it touched.Analysts, audits, the LLM (its own history)
Eval storeWhether the agent did it well. Golden-set versions; eval runs per agent version with per-case scores, traces and the release verdict; and in-chat feedback (rating, flag, correction) linked to the action-log entries behind each answer.Agent release gate, eval reviewers, monitoring

LLM interface

One LLM interface serves every environment. It reads experiment results from the metadata store, starts jobs and checks their status, all through tools in the tool registry. The approach is guardrails, not handcuffs: the LLM gets real access, and the guardrails shape how it uses it.

Capabilities and guardrails

CapabilityWhat the LLM doesGuardrail
Read resultsLists experiments and versions, loads summary components, compares and plots themRuns freely; every answer cites model ID, version and environment
Check job statusReads run state from the job store, and step-level status, failure reasons and logs from the workflow engineRuns freely
Start jobsLaunches builds and scoring runs through the tool registry, which hands them to orchestrationSame workflows and checks as the operator UI; explicit target environment; trigger level per environment (e.g. confirm before launching in prod)
Propose promotionAssembles gains, lift, deciles, holdout results and the audit, and requests a stage changeOnly an analyst's signed-in approval changes the stage
Every action is written to the action log. Access rules and trigger levels live in the tool registry and are enforced when a tool is executed. Every answer carries feedback controls, and every agent version passes the eval harness before release (see Agent evaluation).

Reading results from experiments

1
Question
"Compare the top features across the last three versions."
2
Skill loads
Maps the request to read-path actions and the call shape.
3
Tool registry
Agent calls execute_tool; the registry checks access and routes to the read path.
4
Read path
Reads the metadata index, then only the requested keys.
5
Interpreter
Builds the comparison table or plot in the agent's workspace.
6
Answer
Results with figures inline, citing model ID, version and environment.
Read-path actionUsed for
get_detailsDiscover the action catalog instead of guessing
list_models / search_modelsFind experiments by name (fuzzy)
list_model_versionsEnumerate the runs of one experiment
get_model_summaryCompact first look: version, sizes, top features
list_available_componentsDiscover which keys a version has
load_specific_componentsFetch only what's needed: ROC, gains, deciles, coefficients
load_model_metadataEverything; last resort
Design principles
  • Each tool does one job. Reporting tools read; job tools start and check runs; stage changes go through promotion.
  • Keyed component reads. The agent reads the index to discover what exists, then fetches only the keys it needs, keeping context small and answers fast.
  • Skills hold the contract. Each skill maps common requests to actions and documents the call and response shape.
  • The registry holds the schema. Tool definitions carry parameter and return schemas and a version, so skills and tools can be checked against each other.
  • Two status levels. The registry reports whether routing worked; the tool reports whether the action worked.
  • Compare across versions. Because a key means the same thing in every version, a comparison is: list versions, read the same key from each, tabulate or plot.
  • Monitoring as a tool. Drift results reach the agent through the tool registry, the same way metadata does.
  • Taxonomy on demand. The agent reads the taxonomy as markdown, searches it through a hierarchical vector index, or queries it as an MCP server, so field meanings are looked up rather than guessed.

Agent evaluation

The action log records what the agent did. Evaluation records whether it did it well. There are two loops, and they share one golden set and one eval store. A pre-release harness gates every change to the agent, and feedback controls in chat catch what the harness missed.

Before release: eval harness
Replays a versioned golden set of tasks and retrieval queries against a candidate agent version in dev or staging, scores each case, and compares the scores with the current release. Any change to the agent triggers it: the LLM model, prompts, skills, playbooks, tool schemas, hooks, routing rules, the taxonomy, or a rebuilt vector index.
In use: feedback in chat
Every agent answer has controls to rate it, flag it or correct it. Each signal is linked to the trace behind the answer: its action-log entries, skill, agent version and environment. A bad answer can then be replayed, diagnosed and turned into a test case.

What gets evaluated

AreaWhat's measuredHow it's scored
RetrievalTaxonomy lookups and skill selection return the right fields and skills. Measured as recall@k and the rank of the first correct result on labeled queries, including whether near-target fields are flagged.Deterministic, against labeled relevant sets
Tool use and trajectoryThe right tool, parameters and environment at each step, and the right sequence overall: no skipped verification steps and no redundant calls. Confirmation is asked wherever the trigger level requires it. A fluent answer that skipped a check counts as a worse failure than a visible error.Deterministic, against expected call traces; required steps must appear, in order
GroundingNo hallucinated numbers, fields or versions: every number in an answer matches a component the agent read, and every answer cites model ID, version and environment.Numbers checked deterministically against tool results; an LLM judge checks for unsupported claims
Response qualityThe answer is clear, complete and useful for the question asked. It states its assumptions, and it asks rather than guesses when a request is ambiguous.LLM judge against a rubric, calibrated on analyst ratings
Task completionPlaybook tasks reach the expected end state: plan drafted, gates respected, the right job started.State assertions in a dev sandbox, plus a rubric
GuardrailsThe agent ignores instructions planted in content it reads. It never approves its own promotion, and it never launches in prod without confirmation.Every case must pass; one failure blocks release
Cost and latencyTokens, tool calls and time per task.Compared with the current release; regressions are flagged but don't block by default

Before release: an agent version is promoted like a model

1
Change
A pull request changes a prompt, skill, playbook, tool schema, hook, routing rule, the LLM model, or the vector index.
2
Pin the version
Agent version = LLM model + prompts + skills + tool schemas + hooks + routing rules + index version.
3
Retrieval first
Retrieval cases run before the full suite, so a retrieval regression is caught where it starts.
4
Full suite
CI runs the harness on the pull request, replaying the golden set in dev or staging against fixture data. It never writes to prod.
5
Gate check
Scores are checked against thresholds and against the current release. Every guardrail case must pass.
6
Sign-off
A reviewer approves the release under their own identity. The previous version stays pinned and ready to revert to.

Responsibilities split the same way as in model promotion: policy (the thresholds) is versioned in code, evidence (the eval run) is stored on the version, and a person decides.

In use: feedback in chat

Controls on every answer
LLM agent
Top-decile lift is 3.1 for v2026-09-30 against 2.7 for production v2026-08-01 (acme_lookalike, shared registry).
👍 Helpful👎 Not helpful⚑ Flag✎ Correct
Feedback is a signed-in action under the analyst's identity, like an approval. The agent never sends or edits it.
SignalCaptured as
Helpful / not helpfulOne click on any answer
FlagA reason: wrong number, wrong field, wrong tool, should have asked, unsafe action
CorrectThe right value or field, in the analyst's words
ImplicitPlan rejected, confirmation cancelled, question asked again
Each signal is stored in the eval store with the trace ID, agent version, skill and environment.

Closing the loop

Answer→Feedback→Triage→Reviewed case→Golden set→Next pre-release run
Triage
Flags and corrections go to a review queue, grouped by skill. A reviewer confirms the expected answer, labels the failure area (retrieval, tool use, grounding, response quality or guardrail), and adds the case to the golden set. Feedback never enters the golden set without review.
Root cause: fix the harness first
Failures are also grouped by cause: a missing tool, a vague rule or skill, a missing guardrail or hook, or noisy context. Most agent failures are configuration failures, so the fix starts with the skill, tool, hook or context, and the LLM model is changed only when those don't help. Each fix is then checked against the full golden set, so it can't break something else.
Online monitoring
Each week, a sample of production traces is scored by the same judges used before release. Helpful rate, flag rate and judge scores per skill appear on the monitoring dashboard next to model drift, and a drop raises an alert.

Rules

  • Golden sets are code. They are versioned, reviewed, and owned per skill or playbook, and every case is labeled with its area.
  • Deterministic checks first, judges second. Anything that can be checked exactly, such as numbers, tool calls or retrieved fields, is checked exactly. LLM judges handle only what can't be.
  • Judges are checked against people. Each LLM judge is calibrated on human-labeled cases, and re-checked whenever its model or rubric changes.
  • Scores belong to an exact version. Changing any part of the agent version makes its scores stale until the suite runs again.
  • Eval data follows the LLM data rules. Fixtures and captured traces carry no raw personal data.

Orchestration patterns

Orchestration receives job requests from the tool registry, the operator UI or arriving data, and runs them on a managed workflow engine and managed compute. The platform writes the workflow definitions and adds two checks specific to its data: input checks at the start and a completeness check at the end. Everything else (scheduling, provisioning, retries, cleanup, run state and events) belongs to the managed services.

Workflows as definitions
Each pipeline is a DAG or state machine in version control, with its trigger: arriving data, a tool call, a schedule or an operator. Adding a pipeline means adding one definition.
Input checks first
Every build and scoring workflow runs the input checks as its first step. A failure stops the run, and the results are stored on the job.
Managed compute
Steps run as containers or Spark jobs on a managed service, which provisions capacity, retries failed tasks, enforces timeouts and scales back down. Images are built and tested in CI, so a bad image fails in CI, not in a run.
Completeness check last
The final step compares outputs to sources. A run succeeds only if every expected output exists, because the compute service knows a task exited, not that its outputs are complete.
Run state from the engine
The engine owns run state. Its state-change events update the job store, which holds the platform's view of each run (model version, input checks, expected counts, delivery), so the LLM and the operator UI don't depend on any one engine.
Alerts from engine events
State changes, failures and timeouts from the engine and the compute service are routed to notifications, with no alerting code in the steps.
Runs tagged by run ID
The engine's run ID is passed to every step and applied as a tag on the compute it uses, so a run's resources, logs and costs can always be found.
Scoring resolves "production"
Scoring workflows ask the registry for the production version at dispatch and record that version on the job.

Engine and compute options

OptionRoleGood fit when
Airflow
Amazon MWAA
Workflow engine: DAGs, schedules, data-aware triggers, retries, run historyThe team knows Airflow, or workflows span systems beyond AWS
AWS Step FunctionsWorkflow engine: state machines with native Batch, EMR and SageMaker integrationsAWS-native, event-driven workflows with little to operate
AWS BatchCompute: job queues, array jobs, retries, timeouts, state-change eventsContainer-based training and scoring tasks
Amazon EMR
or EMR Serverless
Compute: managed SparkLarge, Spark-shaped data processing
Any engine plus any compute option gives the same system. Choose by the team's skills and the shape of the workloads.

Built as code

The whole system is version controlled, not just the models. Infrastructure is defined in code (for example with the AWS CDK), so every service, store, queue, role and trigger has a reviewed, versioned definition, and an environment is that definition deployed with its own config.

Infrastructure as code
Services, stores, queues, schedules, alert streams and access roles are all declared in code. Nothing is created by hand, so any environment can be rebuilt from the repository.
One definition, many environments
Dev, staging and prod are deployments of the same stacks. Only config differs, which is what makes "environment comes from config" true in practice.
One owner per resource
Each table, bucket and queue is owned by exactly one stack, which also owns its schema. Other layers use it through the platform client, never by creating it.
Packages released by tag
Shared libraries publish to a private package index when a version tag is pushed, and only if the tag matches the package version.
No stored deploy keys
Build pipelines authenticate with short-lived, federated identity (e.g. OIDC) scoped to publishing or deploying.
Changes are reviewed and reversible
Infrastructure changes go through the same review as code, show a diff before they deploy, and can be rolled back by deploying the previous version.

What lives in version control

ArtifactDefinesReleased as
Infrastructure stacksPlatform services, orchestration, read path, monitoring, LLM interface hosting, roles and triggersDeployed per environment, or once for shared resources
Shared librariesAgent framework, ML framework, platform client, LLM knowledgeTagged package versions
Workflow definitionsDAGs or state machines, their triggers, retry and timeout policies, and compute settingsDeployed per environment with orchestration
Tool definitions and access rulesTool schemas, execution strategy, and trigger levels per environmentDeployed to the tool registry
Product-line code and gatesModel configurations and promotion gatesVersioned with the product line
Agent playbooks and skillsStaged workflows, approval gates, skill contractsReviewed and merged; one copy
Golden sets and eval thresholdsGolden tasks, labeled retrieval queries, judge rubrics, and release thresholds for the agentVersioned with the playbooks and skills they test; run in CI on every change to the agent

Models are the one thing deliberately not deployed from version control: they are data, versioned in the registry and moved by promotion. Their configs, gates and evidence are still traceable to versioned code.

Standards alignment

There is no single MLOps standard, and no formal standard yet for LLMs inside MLOps. There are widely used maturity models, governance frameworks, an emerging agent-standards effort, and de facto tooling. This tab maps the architecture to each.

Maturity models and governance frameworks

FrameworkHow this architecture aligns
Google MLOps levels
0 manual · 1 pipeline automation · 2 CI/CD automation
Level 1–2. Pipelines are configuration on the ML framework and run by orchestration; packages and infrastructure are released through CI/CD (see Built as code).
Microsoft MLOps maturity
levels 0–4
Level 3: automated training, registry, lineage and deployment. Level 4's fully automatic promotion is deliberately not adopted: an analyst approves every promotion.
AWS MLOps phases
initial · repeatable · reliable · scalable
Repeatable and reliable: infrastructure as code, separate environments from one definition, a shared registry, gated promotion and managed orchestration.
NIST AI RMF
Govern · Map · Measure · Manage
Govern: access rules, trigger levels and approval policy, all in version control. Map: product lines with their own gates and audiences. Measure: input checks, eval evidence, drift monitoring, and agent evaluation before release and in use. Manage: promotion, re-promoting a previous version, and re-scoring.
ISO/IEC 42001
AI management system
Defined roles (the analyst approves), retained records (approvals on the model version, the action log), and a review loop (monitoring informs retraining decisions).
NIST AI Agent Standards Initiative
launched February 2026
Its first focus, agent identity, authorization and auditing, maps to three parts of the design: approvals made under the analyst's own identity, access rules enforced in the tool registry, and the action log. Treating read content as data, never instructions, addresses its prompt-injection concern.
Vibe coding to agentic engineering
Google, The New SDLC With Vibe Coding, 2026
The agentic-engineering end of the spectrum. Tests cover the deterministic parts (input checks, gates); evals cover trajectory, tool use and response quality. The harness (skills, tools, hooks, routing, context) is versioned, reviewed and evaluated in CI, and people own the architecture and the approvals.
Human oversight
in-the-loop · on-the-loop
Human-in-the-loop for promotion: nothing changes stage without an analyst. Human-on-the-loop for runs: confirmation before prod launches, alerts and monitoring.

Tooling and protocol interoperability

Layer or storeMaps to
Model registryMLflow Model Registry (versions, aliases such as "production", tags) or SageMaker Model Registry (model package versions with approval status such as pending manual approval / approved). The analyst approval maps directly to SageMaker's manual approval.
Metadata storeMLflow run metrics, parameters and artifacts; SageMaker model metadata and model cards. Keyed components map to logged artifacts addressed by path.
OrchestrationWorkflow engine: Airflow (Amazon MWAA), AWS Step Functions or SageMaker Pipelines. Compute: AWS Batch, Amazon EMR (or EMR Serverless) for Spark, or SageMaker training and processing jobs. State-change events through EventBridge or Airflow callbacks.
Job storeMLflow runs or SageMaker Pipelines executions can hold the same record; either way, it carries the engine's run ID.
Tool registryModel Context Protocol. Tool definitions are published in MCP schema format, and the registry exposes an MCP-compatible JSON-RPC 2.0 endpoint for listing and calling tools, so any MCP client can use the platform's tools.
Read pathThe same role as MLflow's MCP server, which exposes the registry and experiment history to agents; either can sit behind the tool registry.
Agent framework and LLM knowledgeMCP client support for external tool servers; skills as markdown with front matter; the taxonomy served as markdown, a vector index or an MCP server.
Agent evaluationMLflow GenAI evaluation and tracing, or Amazon Bedrock evaluations (including RAG evaluation); Ragas-style metrics for retrieval and grounding; OpenTelemetry GenAI semantic conventions for traces, so feedback and scores join on the same trace ID.
Built as codeAWS CDK (CloudFormation underneath); the same principles apply with Terraform or Pulumi.

Where the design goes further

Most published patterns connect agents to a model registry read-only. Here the LLM reads results, starts jobs and proposes promotions, with guardrails sized to each action: free reads, confirmed launches, and analyst-approved stage changes, all recorded in the action log. The agent is also held to the same standard as the models: each agent version is evaluated and signed off before release, and scored by its users once it is in use.

References

Decisions to make

Choices the design leaves to the implementer. Each has a reasonable default.

DecisionOptions and default
Gates per product lineWhich metrics and thresholds a candidate must pass before the analyst sees it. Default: start from the metrics analysts already review (gains, lift, deciles, holdout gap) and the audit flags.
LLM access controls and trigger levelsWhich actions run directly, which ask for confirmation and which need approval, per environment. Default: reads and status checks run directly; starting jobs asks for confirmation in prod; promotion needs analyst approval.
LLM data and safety rulesWhat data may be sent to the LLM provider, and how content read from files and metadata is treated. Default: no raw personal data in prompts; content the agent reads is treated as data, never as instructions.
Agent release thresholdsWhat an agent version must score before release. Default: guardrail cases must all pass; retrieval, trajectory, grounding and response-quality scores may not fall below the current release by more than a set margin; cost and latency regressions are flagged but don't block.
Golden set ownershipWho writes and maintains the eval cases. Default: each skill or playbook owner owns its cases; a starting set is seeded from the analysts' most common requests and from known failures.
Feedback triageWho reviews flags and corrections, and how quickly. Default: weekly review by the skill owner; guardrail flags are reviewed the same day.
LLM judgesWhere a judge is used, and which model it runs on. Default: only for checks that can't be done exactly; a different model from the agent's; calibrated on human labels before use.
Input checksWhat is checked before training and scoring. Default: schema against the data dictionary, row counts against expectations, and null rates and distributions against a reference.
Retraining triggerWhat prompts a new candidate. Default: an analyst decides from the drift dashboard; the agent can propose.
Outcome feedbackHow real outcomes, when they arrive, flow back. Default: recorded against the model version in the metadata store and shown in monitoring.
Output deliveryWho owns getting outputs to consumers. Default: orchestration delivers and records the delivery on the job.
Deploy automationWhether infrastructure deploys run by hand from reviewed code or automatically on merge. Default: automatic to dev and staging; prod deploys after review of the infrastructure diff.
Retention and backupHow long versions are kept, and how the shared registry is protected. Default: keep every version that was ever in production; back up the registry and its artifacts.
Cost limitsCaps on compute and LLM use. Default: a maximum capacity per compute environment or queue, a per-run concurrency limit, and a per-session LLM budget, with alerts.
Model routingWhich LLM model each step uses. Default: the most capable model for planning and analysis; smaller, cheaper models for routine steps such as taxonomy lookups, summarizing tool output and simple status questions; judges on a different model from the agent's. Routing rules are part of the agent version, so changing them re-runs the eval.
Workflow engine and computeWhich managed services run orchestration. Any of them gives the same system: Airflow (MWAA) or Step Functions as the engine, and AWS Batch, EMR or SageMaker jobs as compute. Default: choose by team skills and workload shape, such as Airflow where the team already runs it, Step Functions for AWS-native event-driven work, Batch for container tasks, and EMR for Spark. Whatever is chosen, keep the input checks, the completeness check and the job store, so the LLM tools stay the same.
Online servingBatch only, or an optional real-time serving layer. Default: batch only until a product line needs real-time scoring.