RAG-CM — the cost model: entry points and how to use them¶
RAG-Stack has one production cost-model package: rag_stack.cost_model.
Its entry class, rag_stack.cost_model.assembly.RAGCMAssembly, predicts the
serving
performance (qps / latency) of a resolved RAG pipeline on declared hardware
from analytical component models plus a closed-loop discrete-event simulator,
with no GPUs at pricing time.
Uncalibrated hardware-spec mode¶
Strict calibration remains the default. For a newly declared hardware system whose calibration campaign is incomplete, opt into the special exploratory mode in the run config:
The equivalent optimize override is:
python -m rag_stack optimize \
--config configs/rag_stack/sgs_a100x8_dragonball_agent_s45.yaml \
--cm-calibration-mode uncalibrated_spec
The CLI persists the value in the project's rag_stack_config.yaml; resume and
hardware replay therefore retain the same policy. Use
--cm-calibration-mode strict to switch back explicitly.
The component rule is exact-profile-or-spec:
- if the exact component profile exists and validates, the CM uses it;
- if it is genuinely absent, the CM uses that component's analytical hardware-spec path instead;
- a present but malformed, stale, ambiguous, or identity-mismatched profile still fails closed;
- profiles are never borrowed across hardware, model, precision, or component identity.
GenZ-backed components use the declared GPU roofline with unit efficiency and no calibrated intercept. FAISS uses the declared CPU/cache/bandwidth facts. LLMLingua2 and multi-list RRF have no complete wall-service model derivable from a datasheet, so their special-mode paths are explicitly labelled lower bounds; host/Python wrapper work remains uncalibrated.
calibration_profiles/stage_batch follows a stronger all-or-nothing
fairness rule. Special mode enables stage overhead only when the exact
system_hardware file contains one complete, valid six-profile generation:
generator
query_expansion
generator_prefill
generator_decode
query_expansion_prefill
query_expansion_decode
If the file is absent, contains only the valid collocated pair, or lacks any of the four direct P/D roles, every stage-batch curve is disabled for every candidate in that run. A malformed file still raises. The decision is frozen once per run and propagated unchanged to DES workers, so collocated and disaggregated candidates cannot receive different calibration coverage.
Sweep output records the policy in cm_calibration_mode,
stage_batch_profile_enabled, stage_batch_profile_reason, and
stage_batch_profile_generation_sha256. Special-mode predictions are for
new-hardware exploration, not calibrated benchmark publication. The full
contract and fallback details are in the
calibration handbook.
The frozen input contract¶
The CM accepts exactly two data inputs, plus configs:
- Workflow schema (
rag_stack.rag_ir.WorkflowSchema) — the STRUCTURE: dataflow type (sequential/agentic), which performance stages exist, the corpus-level statistics (corpus_stats, see below), and — for the no-trace static path — the aggregateruntime_stats. - Quality trace (canonical envelope) — the DATA: one complete invocation DAG per dataset question, produced by the quality run itself. Tokens and bytes only; no runtime observations (no padding shapes, batch sizes, latencies, QPS, timestamps).
The pricing path never reads RAG E2E performance.json, measured serving
records, or benchmark GT. It does read serialized calibration profiles derived
from standalone component measurements; those profiles are CM implementation
artifacts, not protocol inputs.
Quality trace envelope¶
{
"queries": [
{
"question_id": "<unique, non-empty — one per dataset question>",
"invocation_id": "<optional: which runtime invocation won this row>",
"calls": [
{"stage": "semantic_retrieval_encode", "input_tokens": 7,
"output_tokens": 0, "step_idx": 0, "model_id": "mpnet",
"input_bytes": 28, "output_bytes": 0},
{"stage": "generator", "input_tokens": 2048, "output_tokens": 22,
"step_idx": 1, "model_id": "Qwen/Qwen2.5-7B-Instruct"}
]
}
],
"provenance": {"lossy": false}
}
Envelope rules enforced by rag_stack.rag_ir.validate_quality_trace_envelope
(at the producer and again at CM intake) are:
- every included
question_idis unique and non-empty; - Call whitelist (hard-rejected if anything else appears):
required
stage,input_tokens,output_tokens,step_idx; optionalmodel_id,input_bytes,output_bytes. - Bytes are optional, never zero-filled — an absent byte count means
"unknown" (token-based fallback);
0means "zero bytes". - Every query must be complete: its final call is the terminal
generatorcall. provenanceis free-form and cannot affect pricing; selected fields may be copied into diagnostics (the legacy converter records lossy status there).
Dataset cardinality, row order, and provenance.lossy=false are additional
producer/replay admissibility checks because the generic envelope validator
does not know the source dataset. A structurally valid canonical envelope can
therefore be lossy; paper replay must reject it unless its frozen manifest
establishes the stronger dataset contract.
Anything that is not the canonical envelope is rejected with an error pointing at the one-off converter (see "Legacy artifacts" below).
Corpus statistics (WorkflowSchema.corpus_stats)¶
Corpus-level distributions the per-question trace cannot carry — they ride on the schema (never as config side channels):
| Field | Meaning |
|---|---|
avg_chunk_tokens, avg_query_tokens, avg_output_tokens, prompt_template_tokens |
base token shape (required) |
bytes_per_token |
byte/token ratio for communication volumes |
chunk_token_quantiles |
chunk-length distribution grid — powers the analytic reranker-padding price (E[max over a pad group]) |
avg_compressor_elems_per_chunk, avg_compressor_elem_tokens |
compressor element shape |
avg_*_input_tokens overrides |
optional per-stage refinements |
vectordb_index_sizes |
vectordb name → N vectors (FAISS pricing; missing N is an error, never a silent default) |
build_workflow_schema(pipeline_config) lifts them automatically from the
producer-stamped config; CorpusStats.token_stats_dict() reproduces the
original stats dict byte-for-byte.
The three public entry points¶
from rag_stack.types import SystemConfig
from rag_stack.cost_model.assembly import RAGCMAssembly
hardware = SystemConfig.from_dict(system_dict)
assembly = RAGCMAssembly(hardware)
# 1. sweep + trace (the optimizer's normal dynamic path):
# CM sweeps the system space internally and prices each deployment
# against the recorded quality workload.
assembly.evaluate_dynamic(pipeline_config, workflow_schema=schema, trace=envelope)
# 2. sweep, no trace (CM-init / static replays / rag_ir_mode=static):
# aggregate information rides on schema.runtime_stats.
assembly.evaluate_static(pipeline_config, workflow_schema=schema)
# 3. fixed deployment (THE way to test the CM):
# price exactly one deployment — no sweep.
assembly.evaluate_fixed(pipeline_config, system_config, trace=envelope)
All three return an AssemblyResult: performance_score (the selected
scalar, e.g. max-throughput qps), rago_result_df (per-deployment rows:
qps, latency_s, per-stage prices), best_hw_params.
Fixed mode — minimal working example¶
Every evaluated trial archives the exact triple the fixed mode needs:
import json
from rag_stack.types import SystemConfig
from rag_stack.system_layout import cm_fixed_pipeline_config
from rag_stack.cost_model.assembly import RAGCMAssembly
d = "tmp_outputs/rag_stack_projects/<project>/evaluations/eval_0009"
pcfg = json.load(open(f"{d}/pipeline_config.json")) # resolved pipeline (carries _token_stats)
scfg = json.load(open(f"{d}/system_config_resolved.json")) # resolved deployment layout
trace = json.load(open(f"{d}/execution_dag.json")) # canonical quality-trace envelope
hw = SystemConfig.from_dict(cm_fixed_pipeline_config(pcfg, scfg)["system"])
result = RAGCMAssembly(hw).evaluate_fixed(pcfg, scfg, trace=trace)
print(result.performance_score, result.rago_result_df.iloc[0]["latency_s"])
system_config is the resolved layout JSON (resource_groups, per-engine
devices/tp/pp, batch_size_request/batch_size_decode) — exactly the file
the runtime archives. The CM derives collocation groups, chip counts and
retrieval servers from it; nothing else is consulted.
Sweep mode — how the optimizer feeds the CM¶
With system.performance_source: cost_model in the run YAML, each trial's
quality evaluation records the per-question execution trace at runtime; the
controller hands (pipeline_config, schema, trace) to evaluate_dynamic,
and the CM sweeps system.cm_search_space internally to select the best
deployment. No measured data is involved anywhere in that loop.
Independent shortlisted deployment signatures are simulated in a persistent
spawn pool whenever more than one candidate exists. The automatic worker count
uses the process CPU affinity minus two CPUs reserved for the parent and host
services (and never exceeds the shortlist size). Use CM_DES_WORKERS=1 (or
CM_DES_SERIAL=1) for a strict serial replay. Setting CM_DES_WORKERS=N can
request a smaller diagnostic pool but remains capped at affinity minus two.
CM fixes OMP, OpenBLAS, MKL, NumExpr, Accelerate and BLIS to one native thread
before spawning the pool, so an optimizer-level CPU count is expressed as
independent DES processes rather than nested sleeping thread teams.
The same affinity-minus-two ceiling is shared by per-stage generation and
RAGO placement-strategy pricing; each phase additionally caps its workers at
the number of independent tasks available in that phase.
Each adaptive DES saturation probe uses two concurrency turnovers for warmup
and two for its deterministic measurement window; the offered concurrency then
grows until real server queues show sustained backlog in at least three of the
five simulated-time samples at 10/30/50/70/90% of the window. Only
_Engine.waiting > 0 or _StageWorker.queue > batch_cap counts;
logical-admission/frontend occupancy does not. QPS is reported only from that
proved window and never participates in saturation detection or fitting. This
bounds token-step work without making client concurrency a workload axis or
changing any queue, batch, kernel, or transfer physics. The public
simulate(..., warmup_turnovers=..., window_turnovers=...) surface remains
available for longer convergence audits. Production decode additionally uses
a lazy logical-token epoch: resident input/generated-token aggregates and
completion buckets replace the former full resident scan on every token while
preserving scheduler-step prices and FCFS completion order exactly. Set
CM_DES_DISABLE_LAZY_DECODE=1 only for an eager-vs-lazy numerical audit.
For trace-driven throughput and mean-latency selection, the legacy
DynamicEngine trace walk is bypassed. RAGO's portable per-stage capacity
surface first forms a coefficient-free, population-free saturated bottleneck
bound. Large sweeps are then shortlisted independently within active-dataflow x generator-PD x
query-expansion-PD strata, preserving multiple exact device/TP/PP topologies,
multiple request/decode batch parents, and every dynamic-timeout sibling of a
retained parent. The closed-loop DES remains authoritative for every retained
row and for the final winner; the analytical layer does not infer queue wait or
occupancy from a chosen client population. All original rows remain in
rago_sweep.csv: pruned rows are
explicitly marked simulator_status=screened_out, with QPS zero and infinite
latency, and selection considers only simulator_status=ok rows.
Ordinary sweeps retain five batch parents for each of four topologies per
stratum plus a global top-ten parent rescue. Cartesian sweeps with at least
4096 normalized rows retain one parent per topology plus the same global
rescue, keeping full DES near O(100) candidates. These are search-compute
budgets, not calibration coefficients: they do not alter any retained
candidate's physical prediction. The applied budget is recorded as
hybrid_parents_per_topology, and every timeout sibling remains atomic.
Fixed mode and small sweeps remain exhaustive. Set CM_DES_EXACT_ALL=1 to
force exhaustive DES for an audit A/B. CM_DYNAMIC_ANALYTICAL_DIAGNOSTICS=1
restores the separate legacy trace walk; selection policies which consume TTFT
or TPOT also retain that legacy path until DES publishes those events. Neither
switch changes the schema+quality-trace protocol, the RAGO design space, or
candidate identities. No new calibration/assembly coefficient is introduced.
RAG-CM requires cm_search_space.no_microbatching: true. This means every
stage before decode shares batch_size_request; it does not tie decode to
that batch. RAGO independently sweeps the generator decode concurrency and
RAG-CM materializes each RAGO batch_size_generator_decode value one-to-one
as the deployed batch_size_decode. Every (request batch, decode batch)
candidate is therefore a distinct configuration. Its identity is always
validated and retained; in a large sweep it is either explicitly screened or
simulated exactly once. Divergent axes or duplicate normalized configurations
are rejected instead of being hidden by signature grouping.
The requested decode grid is CM-owned and is passed into RAGO explicitly; a
component Pareto frontier is not allowed to truncate deployable resident
depths. By default it is the powers of two through
max_batch_size_request. An explicit cm_search_space.batch_size_decode grid
adds those depths while retaining the decode=request baseline, and scalar
system.batch_size_decode pins each row to max(request, pin). Concrete-device
repricing and DES determine performance after the candidate exists.
Legacy artifacts (pre-freeze traces)¶
Old execution_dag.json files (bare lists, or retired envelopes with
trace_scope) are rejected by the CM. Convert them once, offline:
# one file
python benchmarks/rag_cm_accuracy/tools/convert_legacy_traces.py \
path/to/execution_dag.json --scope measured_only --out converted.json
# a whole optimizer project, in place (.v1.bak backups, idempotent)
python benchmarks/rag_cm_accuracy/tools/convert_legacy_traces.py \
--project tmp_outputs/rag_stack_projects/<project> --scope auto
--scope auto (recommended for projects, handles mixed eras) reads each
file's sibling performance.json: performance.trace_scope ==
"dataset_quality_first_completion" → lossless upgrade (question_id = the
dataset row position as str, cardinality checked against
quality_rows_completed); anything else → measured_only.
--scope measured_only treats a bare list as a saturated measured admission
queue (lossy projection: complete rows only, first occurrence per input
shape, synthetic legacy_shape_<i> ids); --scope first_completion upgrades
already-per-question lists losslessly. The provenance block records exactly
how lossy the conversion was, forever. Resume any pre-freeze project only
after converting it — the controller refuses legacy execution_dag.json
with an error naming this converter.
Calibration profiles are not inputs¶
The production CM and calibration CLI resolve one run-global profile root,
calibration_profiles/ by default, with component
profiles in llm_sim/, compressor_llmlingua2/, faiss_ivf/,
faiss_hnsw/, and retrieval_worker/, and stage curves in
stage_batch/<system_hardware>.json.
Those paths are not RAG-config inputs. system.rag_cm accepts only
calibration_mode; per-config component_calibration_store and
calibration_store fields are unsupported. Remove them instead of pointing
them at calibration_profiles. Use
--profile-dir while staging a campaign or the process-wide
RAG_STACK_CALIBRATION_PROFILE_DIR when intentionally activating a
different complete profile root. A candidate or target-hardware YAML must never
select different calibration coverage from another candidate in the sweep.
Migrating to a new system starts with a read-only, isolated inventory:
rag_stack calibrate status --config configs/calibrate/<box>.yaml \
--scope core \
--data-dir /path/to/stage/calibration_data \
--profile-dir /path/to/stage/calibration_profiles
That command audits the core component and generator coverage. For a full
RAG configuration, use --scope pipeline; it expands the actual pipeline and
checks all required component and stage-batch evidence.
The complete identity matrix, standalone-evidence rules, migration order,
source hashes, and accuracy guards are in
RAG-CM calibration and hardware migration. RAG E2E GT,
measured host wall/padding, and generic correction factors are forbidden fit
sources. Fixed134 is a regression guard only. All profiles are read-only during
pricing; build corpus × embedding workload artifacts explicitly with
the complete rag_stack calibrate workload command documented in the
calibration guide before a reproducible production run.