A machine learning pipeline in production moves data through ingestion, validation, feature computation, training, evaluation, deployment, and feedback. It should produce trustworthy artifacts at each stage, despite schema changes, feature variations, delayed labels, or retraining/roll-back of ML/AI models.
The challenge increases as AI-generated pipeline components become more prevalent. Summarisers, extraction agents, generative classifiers, and embedding models can introduce additional dependencies that must be managed throughout the pipeline.
A prompt, tool list, retrieval rule, routing policy, or model parameter can change the pipeline’s inputs, outputs, and failure modes just as much as a classical feature transformation. These components should be versioned, locked, and pinned to the specific trained model they support.
Hence, a powerful machine learning pipeline architecture separates data management from model deployment. The data layer defines and versions model inputs, while the runtime layer controls which model artifact is active in production. The model artifact forms the contract between these layers, recording the exact versions of the datasets, features, prompts, and configuration used during training.
The article explains the interactions among components, data flow, and control points in such a machine learning pipeline architecture, necessary to optimize AI-supported ML pipelines.
Machine learning pipeline architecture overview
A production machine learning pipeline can be viewed as having two distinct layers.
- The data layer includes everything that produces or defines the inputs a model consumes.
- The runtime layer is where the release actually happens.
In this architecture, LaunchDarkly AgentControl versions the prompts, models, parameters, and tools used to compute AI-generated input features. It manages those feature definitions; it does not replace pipeline orchestration, feature storage, model training, or model-serving infrastructure.
Layer | What lives here | Cadence | What changes |
|---|---|---|---|
Data | Datasets, feature stores, classical feature definitions, AgentControl configs for AI-generated features | Training cycles, with lockstep versioning | Versioned inputs that are locked and consumed at training time |
Runtime | LaunchDarkly CodeControl feature flag wrapping the model artifact pointer, plus guarded rollout on that flag | Per release | The active model variation; nothing else "rolls out" |
The model artifact bridges the two layers. It records the exact data-layer versions it was trained against. At runtime, it becomes the object promoted through the model feature flag.
Keeping these layers separate makes it easier to manage architecture changes independently for safer evolution.
- Data-layer changes create new training inputs.
- Runtime-layer changes decide which trained model artifact serves production traffic.
Mixing the two can create training-serving skew, especially when AI-generated features are involved.
The following diagram shows how these layers connect, which versioned artifacts move between stages, and how production telemetry feeds back into validation and retraining.

A production ML pipeline creates versioned artifacts in the development layer, promotes an evaluated model through a controlled runtime rollout, and feeds production telemetry back into validation and retraining.
Runtime discipline is necessary for the architecture to work
The architecture only works if versioning and rollout discipline are enforced consistently through the following best practices.
Version AI-feature definitions in AgentControl
Each AI step, such as a prompt, model, parameter set, or tool list, is a config variation and must be treated as a feature-engineering definition in the data layer. Iterate on it, evaluate it offline, then lock the variation before training.
Train against locked feature-definition versions
The downstream model is limited to the precise versions of AI features that it used for training. Pin those versions into the training manifest alongside the dataset, code, hyperparameters, and definitions of classical features,
Wrap the trained model in a feature flag
The flag variation should hold the model artifact pointer, such as s3://models/fraud/v1.4.2/model.pkl, plus the AI-feature config variation versions the model expects. The model-serving service, the application component that loads model artifacts and handles inference, evaluates the CodeControl feature flag and loads or selects the active artifact.
Promote the candidate model through a guarded rollout
Model quality, operational, and business KPIs are attached as guardrails for shifting traffic between model variations in phases, such as 5%, 25%, 50%, and 100%. If KPIs decline, the deployment is halted or reversed.
In this design, the model feature flag is the only runtime release control. The selected flag variation identifies the deployed model artifact and the AI-generated input-feature definitions that artifact expects.
Never change an AI-generated input-feature definition used by a deployed model without retraining
A new prompt, model, temperature, AgentControl config, feature flag, or tool list changes the feature definition. If the pipeline contains AI-generated features, guarded rollouts become even more important.
Classical transformations are typically designed to be deterministic, whereas generative feature steps can introduce additional output variability. Similar inputs can produce different outputs depending on the underlying model, prompt, or configuration. While these variations may appear minor, they can significantly influence the behaviour of downstream models that rely on them, introducing regressions that are difficult to predict through offline testing alone.
A guarded rollout becomes the fundamental governance control mechanism for maintaining reliability.
Summary of key machine learning pipeline architecture concepts
Concept | What it means |
|---|---|
Pipeline boundaries | The upstream producers, downstream consumers, and interfaces around the ML system. The model feature flag is the runtime boundary because that is where a release actually happens. |
Pipeline orchestration | Keeps data-layer orchestration separate from runtime rollout. Schedulers like Airflow, Dagster, and Prefect manage ingestion, validation, feature computation, training, and evaluation. The model feature flag manages production traffic. |
Data ingestion and validation | Prior to data entering feature pipelines, create specified batch and streaming ingestion paths with validity checks. Raw data should be examined before being used as a reliable model input. |
Feature consistency | Use the same feature definitions for training and inference. AI-generated features are also feature definitions, stored as AgentControl config variations. Once a model is trained against a variation, that variation becomes part of the model’s contract. |
Reproducible training | Version every input to a training run, including every AgentControl config key and variation version used by any stage that computes or consumes an AI-generated input feature. • Dataset • Code • Hyperparameters • Runtime environment • Container digest |
Model promotion | Promote models through a guarded rollout on the model feature flag. The active flag variation is the only thing being rolled out. AI-feature config variations are included as they are recorded in the model artifact’s metadata. |
Model serving | At request time, the model-serving service resolves the active model variation, loads the corresponding artifact, and reads the AI-feature variations the model expects. Application code should not branch on model version. |
Monitoring and runtime governance | Log • Model version • AI-feature variant versions • Feature hashes • Forecasts • Request information • Judgments made downstream • Final labels. Steady-state health should be monitored using the same measures as those used for gate rollout. |
Pipeline boundaries
Pipeline boundaries define where the ML system begins, what it owns, and how its outputs are consumed. Clear boundaries help catch interface failures like schema drift, missing fields, late data, or ownership gaps before they become model behaviour.
Upstream systems
Upstream systems consist of application events, OLTP databases, CDC streams, warehouses, third-party APIs, and external LLM or embedding APIs. These are the raw inputs that pipelines rely on, and are typically not part of the ownership of the ML platform team.
Downstream systems
Inference services, dashboards, internal applications, and automated decision systems are examples of downstream consumers. These consumers rely on the pipeline's outputs to be stable, documented, and versioned.
Consider interfaces to be contracts
Each handoff between the boundaries and the pipeline should be treated as a contract with clearly specified
- Schemas
- Delivery patterns
- Freshness requirements
- Validation guidelines
- Ownership model
Batch tables, streaming topics, feature tables, model artifacts, and prediction APIs all require clear expectations about what they offer and who is responsible for making changes.
The runtime interface is also a contract. The application evaluates the same model feature flag for every request, while the flag’s selected variation may change as a guarded rollout advances or reverts. Traffic can therefore shift between model versions without changing the application call or redeploying the application.
Pipeline orchestration
With the system boundaries defined, orchestration coordinates the work inside the pipeline. Each stage consumes a versioned artifact, produces another, and records metadata needed by the next stage
Every pipeline stage consumes a named artifact, produces a new artifact, and records enough metadata to explain what changed. A validated snapshot of the data, feature tables, AI feature snapshots, training manifests, model artifacts, assessment reports, and deployment requirements are a few examples of artifacts.
Artifact-based handoffs allow the pipeline to be easily rerun and debugged. The system can review the raw input snapshot if validation is unsuccessful. In case of a drop in model quality, the team can compare training manifests, feature versions, and evaluation reports. When a stage must rerun, the stage may begin from the last trusted artifact, rather than from the beginning of the pipeline.
The following simplified Dagster example represents each pipeline stage as an asset. It moves from raw events to validated data, classical and AI-generated input features, a candidate model, an evaluation report, and finally a model-registration payload. Each output carries the artifact identifiers and version metadata required by the next stage.
The important pattern is the explicit artifact handoff. Registration receives an evaluated candidate model plus its feature-definition binding, allowing the runtime release process to promote that artifact without rerunning earlier stages
Data ingestion and validation
Data ingestion and validation define the trust boundary for the rest of the ML pipeline. Before data is used to compute features, train a model, or promote an artifact, the pipeline must verify that it is complete, up to date, correctly formatted, and valid for the point in time it represents.
Separate approaches for historical and live data
Data ingestion must be able to distinguish historical data from live data, as they have different constraints.
Batch ingestion is used for historical training data, backfills, large joins, and aggregate feature computation. It is most concerned with completeness, repeatability, and correctness at a specific point in time. The training set should contain information available when each label was created.
Streaming ingestion supports low-latency feature updates, event-driven inference, and production monitoring. It is concerned with latency, ordering, deduplication, late-arriving events, and a watermarking strategy.
If batch and streaming ingestion are treated as the same problem, compromises will have to be made in both. Batch pipelines are more difficult to replicate, and streaming pipelines are too slow or too reliant on the warehouse-style assumption. A reliable ML architecture plans for both and maintains a consistent feature definition across paths.
Validate before features
Validation should be performed prior to using raw data as model input. Check raw data at ingestion time for schema compatibility, freshness, null values, allowed values, and timestamp validity. These checks detect upstream contract breaks early, before they affect feature tables or training sets.
Once features and labels are joined, conduct another level of validation. These checks can be for distribution shift, label leakage, target imbalance, unexpected joins, and other kinds of transformation failures.
The two validation stages should be treated as blocking validation points. Bad data should not be used in feature computations. If the merged dataset is invalid, it should not be used for training or promotion until the issue is fixed.
This example accepts a raw-event DataFrame and returns the validated DataFrame if its schema and distribution checks pass; otherwise, Pandera raises a validation error before feature computation begins.
Column checks validate types and allowed values, while the DataFrame-level rule demonstrates a distribution check. The 30% purchase-rate threshold is illustrative and should be replaced with a domain-specific baseline.
Feature consistency
Feature consistency means that the values used during training must match the values the model receives during inference. Whether a feature is produced by SQL, Python, a feature store, or an AgentControl config, the definition must be versioned, reused consistently, and pinned to the model artifact that depends on it.
Classical features: define once, serve from one place
In classical ML, the same feature definition should be used for training and inference. A common failure mode is computing training features from a warehouse with full historical context while computing online features from a stream, cache, or approximation.
For example, an offline pipeline might compute rolling_7d_spend using complete historical data, while the online system computes it from a delayed event stream. The feature names may match, but the values diverge, resulting in training-serving skew.
A feature store helps only if it solves both the consistency problem and storage problem. The important requirement is that offline and online paths use the same definition, the same time-window semantics, and the same point-in-time rules.
AI-generated features are feature definitions in AgentControl
In this architecture, an AI-generated model input, such as a summary, structured extraction, classification, or embedding, is an input feature. AgentControl stores the prompt, provider model, parameters, and tools used to compute that feature as a versioned config variation.
The config specifies the feature prompt, model, parameters, and tools to use for computing the feature. All variations are a feature definition with a version. The classical feature rolling_7d_spend can be defined by SQL or Python code, and the AI-generated feature ticket_summary can be defined by an AgentControl configuration variant. They are both producers of things that are used by the downstream model. They both are part of the data layer.
A model trained using AI features will be limited to the exact config variations that yielded those features. If a config change doesn't result in retraining, the model is being served inputs it wasn't trained on. This is the same failure mode as running a classical feature for a while and then changing its SQL.
The process is as follows:
- Keep trying different versions of the config
- Test the config offline
- Freeze the config
- Add it to the training manifest
- Train against the config
- Use this as a fixed config for that model.
For features created by AI, this rule should be clear since they can be easily modified in a UI. The risk is more obvious with classical feature changes as they typically need code review or pipeline modification. AgentControl variations may travel more quickly, hence the need for version locking and training manifests.
This same point-in-time discipline applies to AI-generated features. If a summary, classification, or extraction is used to train a model, the pipeline must record which AgentControl variation produced it and ensure the production model reads the same variation.
This simplified batch example takes event and label DataFrames and returns one training row per label. The key operation is filtering events to event_time < label_time, which prevents future information from leaking into the feature.
The implementation favors clarity over performance; a production pipeline would normally use vectorized, windowed, or feature-store computation while preserving the same point-in-time rule.
Reproducible training
Reproducible training means that a model can be traced back to the exact data, features, labels, code, configuration, runtime environment, and AI-generated feature variations that produced it. The goal is not only to rerun training, but to make every model artifact inspectable, auditable, and safe to compare, promote, or roll back.
Every input pinned
A reproducible training run identifies every input that may affect the model. This means snapshot or source queries for everything stored as data, such as:
- Datasets
- Versions of feature sets
- Label definitions
- Code versions
- Hyperparameters
- Model artifacts
- Versions of the runtime environment
- Flags for deterministic-mode execution
- Versions of the accelerator runtimes, container, and image digest
The AgentControl config key and variation version need to be pinned as well, for any stage that is any stage that computes or consumes an AI-generated input feature.
This version is just like a SQL transform, Python feature function, or feature store definition. For instance, if ticket-summarizer v7 produces summaries used to train a downstream model, that trained model is bound to the input distribution created by v7. Serving it features generated by v8 may create training-serving skew.
Annotate the variations of AI features at the beginning of training and use these as fixed for the lifetime of the model produced. When a change to an AI-generated feature is needed in production, train a new model with the updated variant and promote it via the runtime layer.
Auditable outputs
A training pipeline should not only generate the model binary, but it should also generate metadata and reports. Those outputs should outline the steps taken to construct the model, the data sources, feature definitions, AI feature variations, and the model's performance during evaluation. This record will be used for bugs, promotions, rollbacks, and team handoffs.
A training manifest can record the complete lineage of a candidate model. This example pins the validated dataset, label definition, classical and AI-generated input-feature definitions, training environment, hyperparameters, and resulting artifact.
The exact format matters less than the guarantee: anyone should be able to inspect the manifest and understand what the model was trained against.
Model promotion
Model promotion is the process of moving a validated model artifact into production traffic safely. Instead of redeploying application code, promotion should occur via a guarded change to the model’s feature flag, in which each rollout stage is evaluated against model quality, operational health, and business impact before traffic increases.
Promotion is a guarded rollout of the model flag
Promotion should happen through the model’s feature flag, not through a redeploy. The flag variation contains the model artifact pointer. It is a simple URI like s3://models/fraud/v1.4.2/model.pkl, or a richer payload that also captures the variations of the AI-features that the model supports.
The model-serving service evaluates the flag and selects the artifact referenced by the returned variation. Promotion shifts eligible contexts from the currently deployed model version to the candidate model version; rollback shifts them back.
Run a guarded rollout, for instance across 5%, 25%, 50% and 100% of the traffic. Model quality, operational, and business metrics should all be used to gate each stage. When the candidate rolls back, the rollout is stopped or rolled back.
Throughout the rollout, the model flag remains the runtime release control. Each model variation contains or references the AI-generated input-feature binding that the model expects. As more contexts receive the candidate model variation, those requests also use the candidate’s preconfigured AI-feature definition.
Guardrails on the rollout, not on a separate dashboard
The rollout should be conducted under the same criteria used to assess the health of the production.
Model quality metrics | Accuracy, precision, recall, F1, AUC-ROC, calibration drift, and prediction-distribution drift. |
Operational metrics | Latency (p95 & p99), error rate, throughput, and resource utilization. |
Business metrics | Conversion rate, escalation rate, false-positive rate, click-through rate, accepted recommendations, or any other product-specific KPI. |
If pipelines feature AI-generated features, add the associated AI-step metrics as early warnings. They may include some of the aforementioned scores from the judges, rate of conformance to the schema, latency, or cost. However, the quality and business parameters of the main model should still be the load-bearing guardrails because those parameters correlate to the visit parameters of users.
Rolling back
The deployment of the previous model artifact can be rolled back during rollout; a redeploy is a configuration transition that can be used to undo the configuration change. The audit trail should document who initiated the rollout, what guardrails were added, what events occurred during the process, and why the rollout either failed, stalled, or succeeded. There needs to be a guarded rollout, as there is a disconnect between the offline evaluation and live behaviour.
Generate candidates with agent optimization
For AI-generated feature steps, agent optimization can generate candidate variations against defined metrics and surface a recommended candidate.
That recommendation is an input to the data layer. It is a candidate feature definition, not something to push directly into production. The candidate still needs to be locked, trained against, recorded in the training manifest, and promoted through a guarded rollout on the model flag.
Promoting a change that involves a new AI feature
Changing an AI-feature config in production without retraining is the failure mode. The safe path is to train a new model against the new variation and roll out that model.
A recommended AI-feature definition follows the lifecycle summarized earlier: lock the definition, regenerate derived training features, retrain and evaluate a candidate downstream model, register the candidate and its binding, and then use a guarded rollout on the model flag.
The data layer does not change during the rollout. The runtime layer points at a different model.
Model serving
Choose a serving pattern
The model serving load should be consistent with the load, latency, and feature usage pattern. Batch inference can be used for large-scale scoring, like churn prediction, account prioritization, or nightly risk scores. In practice, online synchronous inference is suitable for the application when an immediate response is required, such as fraud check, recommendation, ranking, and routing.
Asynchronous inference applies when prediction cannot block the request path. The application can enqueue work, return quickly, and have a worker do the prediction later. Hybrid serving consists of pre-computing features and fetching/computing features at request time. It is very normal if certain features are costly to compute and others rely on the currently requested feature.
Agent graphs are used for multi-step AI workflows to specify the versioned graph of handing off between agent-based configs. These graphs are still data-layer definitions. These are not the runtime release surface; instead, they define how features are made.
Stable serving contracts
A serving system needs stable contracts for
- Input and output schema
- Score semantics
- Versioning strategy
- Timeout and retry behavior
- Backward compatibility
- Idempotency for async inference
During a rollout, these contracts become even more important because multiple model versions may serve different request contexts concurrently.
The application should not branch on the model version. The application should call the same serving interface and receive the same documented output schema regardless of which model artifact is selected. If a new model version requires a different input schema, that change should be handled as an explicit schema migration rather than hidden within the rollout.
The serving layer should either translate from the stable external schema to the model-specific internal schema, support a backward-compatible superset schema, or require a coordinated API version change before the model rollout begins.
For classical ML serving, the request path resolves the active model variation from a feature flag, loads the artifact, and serves the prediction. The flag value can change during a rollout, and the serving service picks up the new pointer on the next request.
The following online-inference example assumes that request, features, model_cache, and latency_ms are supplied by the application. It evaluates the model flag for the request context, selects a cached model artifact, generates a prediction, and emits the custom events used by rollout metrics.
When a downstream model consumes an AI-generated input feature, its flag variation returns a metadata object containing the model URI, model version, and AgentControl config key. The application uses the model version as a context attribute so that a preconfigured AgentControl targeting rule returns the definition associated with that model. The model flag remains the rollout control; AgentControl supports input-feature computation.
The following simplified example assumes that default_payload matches this metadata schema and that the application has initialized the other referenced objects.
AgentControl provides the prompt, provider model, and configuration used to generate the input feature while tracking AI metrics. The application converts the result into an embedding and passes it to the selected downstream model.
The binding can use either a targeting rule that matches the model_version context attribute or a separate AgentControl config for each model version. In either case, configure and validate the binding before rollout and keep it unchanged throughout that downstream model version’s lifetime.
The request path keeps the binding intact
At request time, the model-serving service resolves one runtime control point: the model feature flag. The AI-feature variation follows from the active model because the model metadata records the binding.
Use the same randomization unit on the model flag rollout so a given user keeps seeing the same model, and therefore the same AI-feature variation, during a rollout stage.
The model flag is the only release object. Consistency follows from that single control point. Online evaluations on the AI step can monitor schema conformance, judge accuracy, latency, and cost. These are useful early warnings. But the main model’s accuracy and the business KPIs are still the primary deployment guardrails, because those show whether the model is still receiving the feature distribution it was trained to use.
Monitoring and runtime governance
Link production behavior to training and rollout decisions
Monitoring should link the production behaviour to training data, the definition of the behaviour, decisions to roll it out, and the eventual results. Record the version of the model, versions of the AI features, relevant feature values/hashes, the predictions, the metadata of the request, the decisions made downstream, and eventual labels, if available.
For AI-augmented pipelines, the AI-feature variation is particularly crucial. Otherwise, drift analysis will not be able to determine if the regression is from the model or the input data.
Monitor both system and model health
System metrics, like latency, throughput, error rate, and resource usage, indicate the overall health of the service. Model metrics like feature drift, prediction drift, changes in calibration, and performance measures, after labels are received, indicate model behavior.
Measurement criteria that were used to gate the rollout should be continued in a steady state. Continue measuring the same rollout guardrails after the model reaches steady state so that release and production health are evaluated consistently.
If there is an AI step, then online evaluations should continuously sample live AI outputs. Track accuracy, relevance, and toxicity with built-in judges, use custom judges for signals such as schema conformance, and monitor latency and cost alongside those evaluation scores. Configure the online evaluation to score a percentage of actual traffic. A decline in these signals can indicate that the AI-generated feature is drifting from the distribution the downstream model was trained against, particularly after a config change. These metrics can trigger a notification or automated response if the feature drifts significantly. In this architecture, however, an automated response should not independently switch the AI-feature variation pinned to an already-deployed downstream model.
Runtime controls and human-in-the-loop investigation
Most run-time mitigations should be mechanical:
- Pause the rollout
- Decrease the rollout percentage
- Roll back to the previous variation of the model.
These will not be redeployed.
If a guardrail or monitoring alert requires investigation, the responder should first identify the affected model versions and contexts. They can then compare change history and telemetry from the regression window with a healthy period, and compare failing requests or traces with successful ones. Logs should include the model version, AI-feature variation, relevant feature hashes or values, prediction, and downstream decision.
Conclusion
The strongest ML pipeline architectures keep two layers distinct. The data layer produces and versions the inputs a model consumes. The runtime layer has one release surface, the feature flag that wraps the model artifact.
The model artifact is the bridge between the two layers. It is trained against pinned versions of every data-layer input, including AI-feature config variations, and it is the object promoted through a guarded rollout.
Model quality, operational health, and business outcomes become the guardrails that decide whether the rollout advances, pauses, or reverts.
If a prompt, provider model, parameter, or tool change alters an AI-generated input-feature definition, regenerate the affected training features and train and evaluate a new candidate downstream model before promoting it.
For pipelines that depend on generative input features, guarded rollouts should be treated as the default release-safety mechanism. AI-generated features add variability that the downstream model was not trained to absorb unless those variations were pinned, evaluated, and released with the model that expects them.
Where to go next
- AgentControl Quickstart: versioned AI-feature definitions, including prompts, models, parameters, and tools.
- Feature flags: the runtime layer that wraps the model artifact.
- Guarded rollouts: staged traffic shifts with automated rollback against monitored metrics.
- Agent optimization: generate candidate AI-feature variations against defined metrics.
- Metrics in LaunchDarkly: model quality, operational, and business metrics that feed rollout guardrails.
- AgentControl best practices guide: applied patterns for building and operating the system described above.















