RAG-Stack Config Schema¶
This document defines the YAML config format that drives every
python -m rag_stack optimize run. The fully-annotated, copy-pasteable
companion is configs/rag_stack/config_reference.yaml;
this file explains the concepts and syntax rules behind it.
A config is parsed by rag_stack/config_validator.py
(ConfigValidator) and resolved by
rag_stack/performance_context.py +
rag_stack/controller.py. Validation collects
all errors and reports them at once before any evaluation starts.
1. Core concept¶
A config is a search-space specification, not a single pipeline. The optimizer samples points from it and evaluates each on two decoupled objectives:
- quality — LLM-judged / metric ground-truth evaluation of the RAG answers.
- performance — latency/throughput, from one of two sources (cost model or real on-GPU measurement).
The single most important syntax rule follows directly from this:
A list-valued field is a search dimension; a scalar is a fixed constant.
top_k: [1, 2, 4, 8] # the optimizer chooses one of these per trial
top_k: 8 # always 8 — not searched
model: [Qwen/Qwen2.5-3B-Instruct, Qwen/Qwen2.5-7B-Instruct] # 2-way choice
model: Qwen/Qwen2.5-7B-Instruct # fixed
A single-value list (top_k: [8]) is not a dimension either — it is silently
collapsed to the fixed value.
There are three sweep-spec value forms beyond the plain list:
{range: [lo, hi]}— a continuous parameter:{ordered: [...]}— an ordered categorical, declared inline at the knob whose values it orders. String choices are UNORDERED by default (components, index families); a knob with a semantic order that string sort would scramble (model-size ladders:1.5B < 14B < 3B < 7Blexicographically) wraps its values inordered:and the dimension is builtis_ordered=Truein declaration order — all optimizers read the same flag (numeric-kernel treatment + index spacing in the shared search-space ops, SMACOrdinalHyperparameterfor ≥3 values, Ax ordinal encoding, greedy probes the trade-off endpoints first):Numeric lists are auto-ordered already (wrapping one is a harmless no-op);- component: vllm model: ordered: # YAML value order IS the semantic order (size ladder) - Qwen/Qwen2.5-1.5B-Instruct - Qwen/Qwen2.5-3B-Instruct - Qwen/Qwen2.5-7B-Instruct - Qwen/Qwen2.5-14B-Instruct{ordered: <scalar>}and single-valueorderedlists fail the build loudly.- Fields that are inherently lists (e.g.
algo_search_space.vectordb:,node_lines:,modules:,memory_tiers:) are structure, not sweep dims.
2. The six sections¶
The schema is a fixed set of six top-level keys. Order is convention, not enforced. Pre-6-section top-level keys and renamed knobs are hard-rejected by the validator with a message pointing at the current location.
| # | Section | Purpose |
|---|---|---|
| 1 | global |
run-level switches (quality backend, CM engine mode, iteration budget) |
| 2 | dataset |
source QA + corpus parquets |
| 3 | system |
hardware spec + performance source + per-mode deployment search spaces |
| 4 | optimizer |
optimizer type + params |
| 5 | eval_backend_setting |
everything backend-specific except the algo search space |
| 6 | algo_search_space |
the RAG algorithm dims (corpus chunker + vectordb + pipeline) |
3. Section reference¶
1 · global¶
global:
eval_backend: static_gt # static_gt (default) | flashrag — routes the quality evaluator
rag_ir_mode: dynamic # dynamic (default) | static — which CM engine prices a trial
n_iterations: 100 # total GT-call budget (Phase 1 init + Phase 2 MOBO)
eval_backendpicks the quality backend (see §5).rag_ir_modeonly matters underperformance_source: cost_model:dynamic(default) — the DYNAMIC engine: replay each eval's recorded trace + workflow schema (real per-call token counts).static— the STATIC engine: config-driven aggregate RAGO estimate, no trace needed.
2 · dataset¶
dataset:
dataset_name: dragonball_en # REQUIRED — explicit dataset id
qa: datasets/dragonball/qa_en_sample_100.parquet # QA parquet (carries retrieval_gt)
corpus: datasets/dragonball/raw_corpus_en.parquet # corpus parquet
dataset_nameis mandatory — a non-empty string naming the dataset. It keys the IVF cell-imbalance profile (dataset_name__embedding) so the cost model loads the right data-aware scan-count curve. NEVER derived from a file path (the in-project corpus copy has a fixed basename and abs paths differ across machines); two runs over the same corpus must share onedataset_name.- Repo-relative paths. CLI
--qa-data/--corpus-dataoverride these. - If the corpus is raw and the chunker is a search dim
(
algo_search_space.corpus.chunker), it is re-chunked per eval. If it is an already-segmented final corpus, omit the chunker block and it goes straight to the vectordb.
3 · system¶
system:
performance_source: cost_model # cost_model (default) | measured
rag_cm:
calibration_mode: strict # strict (default) | uncalibrated_spec
cpu: { ... } # cost-model CPU spec (cores, flops, memory_tiers, topology)
gpu: A100_80GB_GPU # GenZ preset name, OR a custom {hardware_key, Flops, ...} block
interconnect: # DECLARATIVE comm fabric (cost-model only). Per device-pair,
gpu_gpu: {bandwidth_gbps: 450, latency_us: 0.4, type: nvlink4} # decimal, one-direction GB/s;
gpu_cpu: {bandwidth_gbps: 54, latency_us: 0.8, type: pcie5_host} # no preset table. `type` is a
# FREE-TEXT LABEL only (not resolved to a number). KV handoff
# (prefill→decode) uses gpu_gpu; vectors/text use gpu_cpu.
# Omitted pair / no bandwidth_gbps → gpu.ICN fallback.
performance_objectives:
selection: max_throughput # reduces the RAGO sweep to one perf scalar:
# min_latency (default) | max_throughput |
# min_ttft | min_tpot | constrained
retrieval: # search-time FAISS runtime knobs (BOTH perf modes)
faiss_num_threads: null # int = pin | [ints] = search dim | null/omitted
# # = DERIVE min(batch, physical cores) — the single
# # policy shared by CM and measured.
# faiss_indexing_thread: 30 # BUILD-time OMP threads (kmeans/PQ/HNSW graph).
# # int >= 1 = pin | null/omitted = DEFAULT cpu_count-2.
# # NOT a search dim / NOT in the cost model (build is
# # outside CM measurement); restored to 1 after build
# # so it never pollutes the search-thread anchor.
# faiss_ivf_parallel_mode: 0 # scalar PIN only (0=inter-query, 1=intra-query).
# # CM-managed deployment axis — never an optimizer
# # dim; unpinned, the CM sweeps {0,1} per batch.
cm_search_space: { ... } # cost-model RAGO deployment sweep (see below)
system_design_space: { ... } # optimizer-owned deployment dims (see §4)
gpu/cpu—hardware_keyis the SINGLE calibration locator. A bare preset key (gpu: H100_GPU,cpu: epyc_9124) locates ceff/meff + every calibration profile by itself. A known preset key may NOT also carry explicit specs — it is a hard config error (it would silently shift the base the calibration was fit against). For NEW hardware, give a dict with a freshhardware_key+ all specs (gpu: {hardware_key: my_h200, flops_tflops: …},cpu: {hardware_key: my_epyc, num_cores: …}); the specs flow to GenZ/FAISS and the calibration profiles are saved/located under that new key. (interconnectis exempt from the preset-exclusive rule — it has no calibration profile and no preset table;{type: nvlink4, bandwidth_gbps: 450, latency_us: 0.4}with numbers written directly stays valid,typebeing a documentary label only.)interconnectfeeds both physical communication models.gpu_gpu.bandwidth_gbpsandgpu_gpu.latency_usare physical topology facts. They drive the PD/stage communication model and GenZ's analytical TP/PP model. Production does not load_tp_pp_efficiency, add a fitted NCCL collective residual, or apply a fitted PP penalty. TP and PP remain deployment choices priced from analytical compute, memory and physical communication; they are not calibration-profile identities. Legacy TP/PP fields that remain in an oldcalibration_profiles/llm_sim/<gpu>.jsonartifact are ignored by production pricing and must not be used as calibration evidence.- Calibration profiles are operator-local and use the hardware identity of
the resource that executes the operator:
faiss_ivf/<cpu>.json,faiss_hnsw/<cpu>.json, andretrieval_worker/<cpu>.jsonare CPU-keyed;llm_sim/<gpu>.jsonandcompressor_llmlingua2/<gpu>.jsonare GPU-keyed. A non-hardware workload specialization goes in its named workload namespace (for examplefaiss_ivf/imbalance/<dataset>__<embedding>.json). There is no assembly/E2E/PD scalar profile and nosystem.assemblycoefficient override: placement, communication, batching, and PD behavior are simulated from the declared topology and component service models. - The production CM implementation is
rag_stack.cost_model. All serialized CM inputs and outputs use one canonical structure, validated by exact required fields, types, identities and hashes; artifacts carry no schema-version field and there is no version negotiation. Deployable calibration artifacts live under the canonicalcalibration_profilesroot. system.rag_cm.calibration_modecontrols missing-profile policy for the cost model. Omit it, or setstrict, for the production default: every exact calibration required by a candidate must exist and validate.uncalibrated_specis an explicit new-hardware mode. An exact valid component profile is still used when present; if the exact component profile is absent, the component uses its analytical physical-hardware specification instead. It never borrows a profile from another hardware/model identity and never treats malformed, stale, or identity-mismatched evidence as merely absent. Empirical-only LLMLingua2 and multi-list RRF wrappers use labelled hardware-roofline lower bounds when absent; their Python/host wrapper cost is not implied by the hardware specification. Treat this mode as exploratory, not calibrated evidence.
This is the only supported key in the system.rag_cm block.
Per-config component_calibration_store and calibration_store fields are
unsupported and fail validation. Remove them instead of pointing them at a
directory. Calibration uses the canonical
calibration_profiles root, with an optional process-wide
RAG_STACK_CALIBRATION_PROFILE_DIR override for an isolated campaign.
calibration_profiles/stage_batch has a stronger fairness rule in
uncalibrated_spec mode. The stage correction is enabled only when the exact
hardware file contains one valid complete six-profile generation
(generator, query_expansion, and both direct P/D roles for each). If that
generation is absent or incomplete, none of its stage-batch curves are
used for any deployment candidate. This prevents a collocated candidate from
receiving a calibrated correction while a disaggregated candidate is priced
without the corresponding correction.
The optimize CLI may override and persist this field in the project config:
python -m rag_stack optimize \
--config configs/rag_stack/sgs_a100x8_dragonball_agent_s45.yaml \
--cm-calibration-mode uncalibrated_spec
Passing --cm-calibration-mode strict explicitly switches a resumed project
back to strict mode. The setting changes CM pricing policy only; measured
performance remains measured.
- cpu.memory_tiers should stay monotonic (L1 > L2 > L3 > DDR bandwidth)
— the roofline model assumes faster upper tiers (not validator-enforced).
Optional cpu.topology block (validated): uniform (default) / numa
(needs sockets × cores_per_socket == num_cores) / heterogeneous
(needs p_cores.count + e_cores.count == num_cores). NUMA and heterogeneous
are mutually exclusive.
- system.retrieval holds search-time retrieval runtime knobs (never part
of the index-build signature). faiss_num_threads: an int pins, a list is an
optimizer search dim, and null (or omitting the key) means DERIVE
min(batch, physical cores) — the one policy shared by CM and measured; an
explicit null just documents that intent in the YAML.
faiss_ivf_parallel_mode is system design space managed by the
cost model — never an optimizer dim: unpinned, the CM prices both modes
per batch inside its own deployment sweep and keeps the best (per-batch
argmin is exactly the global sweep result — the mode consumes no shared
resource); a scalar pins it, a list is rejected (an explicit grid belongs in
cm_search_space.faiss_ivf_parallel_mode). Retrieval uses batch_size_request;
it does not expose a separate batch axis. (Calibration profiles were fit at
1 thread, so faiss_num_threads > 1 is roofline extrapolation in CM mode.)
- cm_search_space — the RAGO deployment sweep used by the cost model:
cm_search_space:
placement_policy: [disaggregated, collocated]
min_num_gpus: 2
max_num_gpus: 4
max_num_gpus_per_stage: 4
gpus_per_server: 4
max_batch_size_request: 256 # positive power of 2
no_microbatching: true # request cohort shared by non-decode stages
allow_decode_collocation: true # may decode share a GPU group with prefill/etc.
cm_search_space.batch_size_decode — decode-engine concurrency
(deployed max_num_seqs), agentic rows only, constraint decode ≥
request. Default: powers of 2 ≤ max_batch_size_request. Pin with the
scalar system.batch_size_decode.
- cm_search_space.dynamic_batch_timeout_s — light-stage dynamic-batching
wait cap. Default [0.002, 0.01, 0.05]. Priced with no effect in the
saturated steady-state model; carried as a deployment-runtime axis. Pin with
system.batching.dynamic_timeout_s.
- cm_search_space.faiss_ivf_parallel_mode — FAISS IVF query-time threading
(0=inter-query, 1=intra-query). Default [0, 1], priced per batch inside
the vector-search stage sweep with the best mode kept (single-mode when no
faiss_ivf index — HNSW ignores the knob). Pin with the scalar
system.retrieval.faiss_ivf_parallel_mode. Never an optimizer dim.
4 · optimizer¶
optimizer:
type: agent_qnehvi # from OPTIMIZER_REGISTRY
params:
seed: 44
n_both_init: 10 # # of joint-objective init points (Sobol) before BO
# ... optimizer-specific params
Registry (rag_stack/optimizer/__init__.py):
type |
What |
|---|---|
sobol |
pure Sobol quasi-random baseline (no surrogate; hierarchy-aware) |
greedy |
budget-aware two-mode greedy baseline (see below) |
ax_qnehvi |
qLogNEHVI via Ax (single-task GP per outcome) |
ax_qnehvi_mixedgp |
qLogNEHVI with a mixed categorical+continuous kernel |
smac_mobo |
SMAC3 ParEGO (random-forest scalarized MOBO) |
agent_qnehvi |
main optimizer — LLM-agent-guided qLogNEHVI |
agent_qnehvi_slo |
agent_qnehvi + suggest-time SLO gating via the static cost model |
greedy_then_hv_anchor_agent_qnehvi |
greedy first pass, then agent_qnehvi with HV-band prompt discipline and expected-HV-gain pool ranking |
Common params:
seed— RNG seed.objectives— list of{objective_name, ref_point, cost}forperformanceandquality(SMAC/Ax HV reference points;costis the decoupled eval-cost ratio).- Init-budget naming differs by optimizer (a known wart):
ax_qnehvi/ax_qnehvi_mixedgp/agent_qnehviusen_both_init.smac_mobousesn_gt_init(aliased ton_both_initinternally).- Agent optimizers always construct their LLM Agent and every post-init BO
round attempts an Agent proposal; this path cannot be configured off.
agent_modelselects the model (defaultdeepseek-v4-flash). composite_objectives.quality.sub_metricsis project/result metadata used for expanded CSV columns. It is auto-derived fromeval_backend_setting.metricswhen descriptions are present and is not an optimizer constructor control. The Agent reads the observed metric payload and the evaluator metric definitions directly.
greedy params (the AutoRAG-style baseline ladder — pure black-box, no CM
side information; the TOTAL GT budget n_iterations is auto-injected from
global.n_iterations / benchmark gt_budget and split across pipeline
stages proportionally to log2(1 + #options), largest-remainder integerized;
a method-level shape switch like rag_dataflow forks into independent
per-branch chains):
optimizer:
type: greedy
params:
search_mode: forward # forward = AutoRAG-faithful single pass |
# lookback (default) = + backward/forward
# refinement passes (geometric budget split)
beam_width: 3 # breeding slots per branch, by ROLE: both front
# ends (span anchors) + interior points by
# exclusive-HV contribution (the knee).
# 3 = fast end + accurate end + knee (default);
# 2 = ends only; 1 = plain single chain
stage_weight: log # log (default) | uniform (allocation ablation)
pass_budget_fraction: 0.6 # γ — lookback only: each pass gets γ of the
# branch's remaining budget
max_passes: 8 # lookback pass cap
# per_stage_cap: 4 # fallback cap, used only when no n_iterations
# budget is injected
5 · eval_backend_setting¶
Everything backend-specific except the algo search space. Keys are
interpreted per global.eval_backend.
static_gt:
eval_backend_setting:
queries_per_trial: 80 # measured-mode subsample (head N of qa.parquet); omit = all
combined_quality: # how sub-metrics fold into the single `quality` objective
mode: mean # mean | weighted (weighted takes {metric: weight})
metrics: [deepeval_answer_correctness]
metrics: # the full metric set (target + auxiliaries)
- metric_name: deepeval_answer_correctness
model: openai/deepseek/deepseek-v4-flash
description: GEval factual correctness vs gold answer. THE business target.
- metric_name: deepeval_context_recall
model: openai/deepseek/deepseek-v4-flash
description: Retrieval ceiling.
- metric_name: retrieval_token_recall # local token-overlap metric (no LLM judge)
combined_qualitydefines the scalarqualityobjective the optimizer maximizes (usually just the business-target metric). All othermetricsare auxiliary signals kept for surrogate reconstruction / stage diagnosis.metricsis the full vocabulary, validated against the backend's metric set. Each LLM-judged metric needsmodel; local metrics (e.g.retrieval_token_*) need none.descriptionis the one recognized meta-field — it feeds the agent optimizer's auto-derivedcomposite_objectives.quality.sub_metrics(§4); all other keys are forwarded to the metric function as kwargs.performance_only: trueis a standalone measured-replay/benchmark switch — it skips ALL quality scoring and is rejected in an optimize config.
flashrag:
eval_backend_setting:
metrics:
- metric_name: em
- metric_name: f1
- metric_name: rouge-l
flashrag: # FlashRAG framework settings
framework: vllm
monitor: full
gpu_memory_utilization: 0.85
max_retrieval_num: 5
model2path: { e5: intfloat/e5-base-v2 }
FlashRAG has its own metric vocabulary (em, sub_em, f1, acc, recall,
precision, bleu, rouge-1/2/l, llm_judge, input_tokens,
retrieval_recall, retrieval_precision) and no prompt_maker /
power-of-two top_k requirement.
6 · algo_search_space¶
The RAG-algorithm dimensions: the corpus chunker and the pipeline. The pipeline shape depends on the backend.
algo_search_space:
corpus: # OPTIONAL — only when sweeping chunking
chunker:
chunk_size: [128, 256, 512, 1024, 2048]
component: [character, recursivecharacter] # any registered chunker (token,
# sentence, sentencewindow, …)
pipeline:
# static_gt → {node_lines: [...]} (sequential only), or
# {rag_dataflow: [...], node_lines: [...]} (sequential+react co-mode)
# rag_dataflow entries may carry the branch's SUB-CONFIG; the declared
# hierarchy is wired into Ax dependents by the search-space builder.
# When react is reachable, its branch MUST declare max_iter — the react
# loop's only round cap:
# rag_dataflow:
# - sequential
# - react:
# max_iter: [3, 5] # REQUIRED on the react branch
# flashrag → {methods: [{mode, node_lines, ...}, ...]}
node_lines:
- node_line_name: retrieve_node_line
nodes:
- stage: semantic_retrieval
top_k: [1, 2, 4, 8, 16] # static_gt: must be powers of 2 (RAGO cost model)
modules:
- component: vectordb
# SCALAR reference — one pinned index; query-time knobs sit flat:
# vectordb: faiss_ivf_index
# nprobe: [8, 16, 32, 64]
# DICT-FORM selector — the index FAMILY is a search dim; each
# family's query-time knob is declared INSIDE its family (ownership
# by nesting, compiled into Ax gating: the non-chosen family's block
# params + query knob are inactive per trial, and the resolver drops
# the unreferenced block from the runtime config):
vectordb:
faiss_ivf_index:
nprobe: [8, 16, 32, 64] # IVF query-time
faiss_hnsw_index:
ef_search: [32, 64, 128] # HNSW query-time
- stage: passage_reranker
optional: true # → the optimizer gets an enable/disable dim
top_k: [1, 2, 4, 8]
modules:
- {component: sentence_transformer_reranker, bits: f32}
- {component: flag_embedding_reranker, bits: f32}
- node_line_name: post_retrieve_node_line
nodes:
- stage: prompt_maker # required for static_gt sequential blocks
modules:
- component: fstring
prompt: "Question: {query}\n Passage: {retrieved_contents}\n Answer:"
- stage: generator # required (≥1 generator)
modules:
- component: vllm
model: [Qwen/Qwen2.5-3B-Instruct, Qwen/Qwen2.5-7B-Instruct]
temperature: {range: [0.0, 1.0]}
Structure & rules:
- A pipeline is a list of node_lines; each node_line has nodes; each
node has modules. User-settable
stagevalues:query_expansion,semantic_retrieval,passage_reranker,passage_filter,passage_compressor,prompt_maker,generator. (hybrid_retrievalis also accepted as a retrieval stage but has no cost-model pricing yet.) optional: trueon a node turns it into an enable/disable search dim (the optimizer may skip it).- Module choice within a node is itself a search dim: listing multiple
modulesmakes the optimizer pick one, and that module's own list-valued params are searched conditionally (hierarchy via Axdependents— inactive branches cost no search budget). - vectordb reference is by
nameintoalgo_search_space.vectordb(§6a). A scalar pins one index. Sweeping the index FAMILY uses the dict-form selectorvectordb: {<family>: {<its query-time knobs>}}— compiled into real Ax gating: the selector gates each family's block params and its nested query knob, so the non-chosen family is inactive per trial (clean GP inputs, no eval budget on inert knobs) and the resolver drops the unreferenced block from the runtime config. A flat multi-family list (vectordb: [fam1, fam2]+ flatnprobe:/ef_search:) is rejected — it leaves the other family's knobs as always-active phantom dims. - Query-time retrieval knobs live on the module (never in the vectordb
block): flat next to a scalar reference, nested per family in the dict-form
selector. Declaring
nprobe/ef_searchinside a vectordb block is not rejected, but it silently becomes part of the index build signature — the same on-disk index would be rebuilt per value instead of served as-is.system.retrieval.faiss_ivf_parallel_modeis IVF-only at runtime but stays an always-active system dim — algo selectors never gate system-space knobs (layering); HNSW simply ignores it. - static_gt requirements (enforced by the validator): ≥1
generator, ≥1 retrieval node,top_ka power of 2 (RAGO cost-model constraint), and ≥1prompt_makerfor sequential blocks — a react block needs none (it builds its own agentic prompt). flashrag drops theprompt_makerand power-of-two requirements. - flashrag pipeline is
{methods: [...]}; each method block has amode(sequential,corag,ircot,flare, …) plus method-specific knobs (max_iter,threshold,look_ahead_steps, …). The optimizer'spipeline.modechoice activates exactly one method's subtree per trial.
6a · algo_search_space.vectordb¶
A list of index definitions, inside algo_search_space (the vectordb
section is algo search space — index family/build knobs move retrieval
quality). The retrieval module references one (scalar) or several (dict-form
selector, §6) of these by name; at resolve time the referenced, scalarified
block(s) are re-exposed at runtime config["vectordb"] for the evaluators,
and unreferenced blocks are dropped.
algo_search_space:
vectordb:
- name: faiss_ivf_index
db_type: faiss_ivf
embedding_model: huggingface_all_mpnet_base_v2 # may be a SWEEP LIST;
collection_name: my_index # embedding_dim is derived per model
nlist_factor: [1, 2, 4, 8] # SHARED: nlist = clamp(round(factor·sqrt(N)), 1, N); N = per-eval chunk count
index_type: # SEARCH DIM (hierarchical): optimizer picks pq or flat;
pq: # each type's exclusive knobs are active ONLY when chosen.
M: [8, 16, 32, 48, 96] # PQ sub-quantizers (must divide EVERY candidate model's dim)
nbits: [4, 8] # bits per code
# dsub: [4, 8] # ALTERNATIVE to raw M: target sub-vector dim D/M;
# # M is derived per model (a divisor of D, capped at 64)
flat: {} # IVF-Flat: exact full vectors, no quantization knobs
# FAISS roofline calibration resolves AUTOMATICALLY and exclusively from
# system.cpu.hardware_key: calibration_profiles/faiss_ivf/<cpu>.json.
- name: faiss_hnsw_index
db_type: faiss_hnsw
embedding_model: huggingface_all_mpnet_base_v2
collection_name: my_hnsw_index
M: [16, 32, 48] # build-time graph degree
ef_construction: [100, 200, 400] # build-time
Rules:
name/db_type/embedding_modelare required.faiss_ivfandfaiss_hnsware the cost-modeled families; other runtimedb_types (chroma, milvus, …) exist but have no cost-model binding, so CM mode rejects them.embedding_modelis a legal search dim (list = the optimizer picks the encoder per trial).embedding_dimis derived from the model for the known fixed-dim local models — omit it there (an explicit value that conflicts with the derived dim is rejected). Only configurable-dim models (openai/mock) still require an explicitembedding_dim.- Only build-time params live here. Query-time knobs (
nprobe,ef_search) go on the retrieval module (§6), so the same on-disk index serves every value without a rebuild. When the module sweeps the index FAMILY, each query knob is declared inside its family in the dict-form selector — ownership by nesting, compiled into Ax gating. - No
pathfor faiss stores. FAISS indexes are content-addressed (chunk-hash + index params) and live in ONE shared global cache (<repo>/.cache/rag_stack/faiss, overrideRAG_STACK_CACHE_DIR), reused across every run/seed — the resolver ignores any per-projectpath, so omit it. (Non-faiss stores like Chroma still declare their project-localpath.) parallel_modeis not allowed here — it is a runtime threading knob (system.retrieval.faiss_ivf_parallel_mode), not an index-build parameter.db_type: faiss_ivfcovers both IVF index types via a nestedindex_typeblock: each key (pq/flat) holds that type's exclusive knobs (pq.Morpq.dsub,pq.nbitsfor IVF-PQ; IVF-Flat has none). The optimizer chooses among the keys, and — because it's a hierarchical (dependent) search dim — only the chosen type's knobs are active per trial (Ax never suggestsM/nbitsfor a flat arm). Shared knobs (nlist_factor, …) stay at the vectordb level. A single key pins that type; a bare scalarindex_type: pqalso works (thenM/nbitssit at top level).- For IVF,
nlist_factorexpressesnlistas a factor ofsqrt(N)so it scales with the per-eval chunk count and can never exceedN(untrainable k-means). A fixednlist: [...]list is also accepted (see the measured baselines). Applies to bothindex_type: pqandflat. ${PROJECT_DIR}(and any${ENV_VAR}) is interpolated at load time;PROJECT_DIRis set to the project root before the config is read.
4. Performance source & system-space owner¶
Two ORTHOGONAL keys (both under system:):
performance_source(omit →cost_model) — WHERE performance numbers come from:cost_model(analytical RAGO/GenZ/FAISS models, no GPUs needed) ormeasured(real vLLM deployment, timed; needsglobal.eval_backend: static_gt).system_space_owner(omit → derived:cost_model→cost_model,measured→optimizer) — WHO searches the deployment design space:cost_model: no optimizer-facing deployment dims; the CM sweeps every deployment internally per trial and selects the best (the decoupled search). Usessystem.cm_search_space.optimizer: the optimizer samples deployment dims fromsystem.system_design_space; exactly ONE deployment is priced (cost_model) or launched (measured) per trial.
Legal matrix: (cost_model, cost_model) = decoupled CM search;
(optimizer, measured) = classic measured mode; (optimizer, cost_model) =
joint-space search priced by the CM (e.g. the RQ1 baseline);
(cost_model, measured) = rejected (a measured run cannot realize a sweep).
Under system_space_owner: optimizer the design space needs the real
available_gpus list (it lives INSIDE system_design_space):
system:
performance_source: measured # or cost_model (joint-space CM pricing)
system_design_space:
batch_size_request: [1, 8, 16, 32] # request cohort for retrieval/prefill/rerank
batch_size_decode: [8, 16, 32, 64] # generator decode continuous-batch cap
dynamic_batch_timeout_s: [0.002, 0.01, 0.05] # non-vLLM dynamic batching wait
vllm_kv_cache_dtype: auto
available_gpus: [cuda:0, cuda:1, cuda:2, cuda:3]
cm_search_space.{min,max}_num_gpus / max_num_gpus_per_stage still bound
the per-trial GPU layout. batch_size_request and vllm_max_num_seqs are
mutually exclusive (the global batch drives the vLLM concurrency knob).
The GPU layout itself is driven by ONE of two mutually exclusive mechanisms:
gpu_layoutcategorical (default): auto-synthesized fromcm_search_spacebounds ×available_gpus; per-engine GPU counts, collocation cuts, TP/PP,gpu_memory_utilizationand serving mode are derived from the partition, not searched. Listing derived knobs (vllm_gpu_memory_utilization,tensor_parallel_vllm,serving_mode, …) insystem_design_spaceis silently ignored (warned); per-stage batch dims (batch_size_generator, …) are a hard error.batch_size_decodeis the only separate batch-like knob (decode-engine continuous batching);dynamic_batch_timeout_sis a global count-or-timeout wait for non-vLLM batched stages.deployment_presetcategorical: named practitioner templates declared undersystem.deployment_presetsand listed insystem_design_space.deployment_preset(suppressesgpu_layout; rawcut_after_*/num_gpus_*/pipeline_parallel_*dims are rejected alongside it):
system:
deployment_presets:
collocated_tp4: {placement: collocated, num_gpus: {generator: 4}}
disagg_2p2d: {placement: pd_disagg,
num_gpus: {generator_prefill: 2, generator_decode: 2}}
disagg_1p2d: {placement: pd_disagg, # asymmetric PD (role-qualified)
num_gpus: {generator_prefill: 1, generator_decode: 2}}
system_design_space:
deployment_preset: [collocated_tp4, disagg_2p2d, disagg_1p2d]
batch_size_request: [16, 64, 256]
available_gpus: [cuda:0, cuda:1, cuda:2, cuda:3]
placement ∈ collocated (every GPU stage in one resource group) |
pd_disagg (decode split off after generator_prefill) |
disaggregated (each sized stage its own group). Under pd_disagg,
auxiliary stages ride the prefill-side group at 1 slot unless sized — an
optional aux: ride | left key controls that placement. num_gpus keys
are engine names or role-qualified stage names (asymmetric PD); values
must be powers of two. TP is derived (chips / pipeline_parallel). A
preset that cannot fit the trial's pipeline shape within the
cm_search_space bounds penalizes that trial (never silently projected).
Quality is deployment-independent, so trials that differ only in system dims
reuse the first trial's judge scores + recorded trace (quality memoization;
every GT eval dir carries a quality_memo.json marker whose hit flag says
whether it was reused).
5. Syntax cheat-sheet¶
| Want | Write |
|---|---|
| Fixed value | scalar: top_k: 8 |
| Discrete choice | list: top_k: [1, 2, 4, 8] |
| Continuous range | temperature: {range: [0.0, 1.0]} |
| Ordered categorical | model: {ordered: [small, medium, large]} |
| Disable-able stage | optional: true on the node |
| Module choice | multiple entries under modules: |
| Choose index family | dict-form selector: vectordb: {ivf_name: {nprobe: [...]}, hnsw_name: {ef_search: [...]}} |
| Env / project path | ${PROJECT_DIR}/data/qa.parquet, ${ANY_ENV_VAR} |
| Quality scalar | eval_backend_setting.combined_quality |
| Stage diagnosis for the agent | per-metric description (feeds auto-derived sub_metrics) |
6. Validation & failure modes¶
ConfigValidator(config).validate() runs before optimization and surfaces all
errors together. Common rejections:
- Any legacy top-level key (
data,pipeline,node_lines,corpus,vectordb,gt_evaluation,n_iterations) → migration error pointing at the 6-section location. algo_search_space.pipelineshape mismatched to the backend (node_linesfor static_gt,methodsfor flashrag), or a react branch withoutmax_iter.- static_gt missing a
generator/ retrieval node, a sequential block missingprompt_maker, or a non-power-of-twotop_k. Mnot dividing a candidate embedding dim (IVF-PQ); anembedding_dimthat conflicts with the model-derived dim; a missingembedding_dimfor a configurable-dim model.- Unknown metric name for the backend.
- A module with no cost-model mapping (CM mode).
system_space_owner: optimizer(e.g. measured mode) withoutsystem_design_space/ itsavailable_gpus.- A flat multi-family
vectordb: [fam1, fam2]module reference (use the dict-form selector).
By default the validator also snapshot_downloads any configured HuggingFace
model that is not cached locally (check_model_cache=True).
7. Minimal static_gt example¶
global:
eval_backend: static_gt
n_iterations: 30
dataset:
dataset_name: dragonball_en
qa: datasets/dragonball/qa_en_sample_100.parquet
corpus: datasets/dragonball/raw_corpus_en.parquet
system:
# Bandwidth has ONE semantic (no memory_efficiency knob): aggregate (all-core
# sustained STREAM) + single-thread (one-core sustained STREAM). Effective
# BW(t)=min(t*single_thread, aggregate); at t=1 (calibrated regime)=single_thread.
cpu: { num_cores: 12, peak_flops: 1.04e12, peak_int_ops: 1.04e12,
mem_bandwidth: 116.5e9, mem_bandwidth_1t: 14.5e9,
memory_tiers: [{name: llc, bandwidth: 200e9, bandwidth_1t: 25e9, capacity: 12e6},
{name: ddr, bandwidth: 116.5e9, bandwidth_1t: 14.5e9, capacity: 36e9}] }
gpu: A100_80GB_GPU
performance_objectives: { selection: max_throughput }
cm_search_space:
placement_policy: [disaggregated, collocated]
min_num_gpus: 2
max_num_gpus: 4
max_num_gpus_per_stage: 4
gpus_per_server: 4
max_batch_size_request: 256
no_microbatching: true
optimizer:
type: agent_qnehvi
params: { seed: 42, n_both_init: 10 }
eval_backend_setting:
combined_quality: { mode: mean, metrics: [deepeval_answer_correctness] }
metrics:
- { metric_name: deepeval_answer_correctness, model: openai/deepseek/deepseek-v4-flash,
description: GEval factual correctness vs gold answer. }
algo_search_space:
vectordb:
- name: faiss_ivf_index
db_type: faiss_ivf
embedding_model: huggingface_all_mpnet_base_v2 # embedding_dim derived (768)
collection_name: my_index
nlist_factor: [1, 2, 4, 8] # shared
index_type: # hierarchical search dim: pq vs flat
pq:
M: [8, 16, 32, 48, 96]
nbits: [4, 8]
flat: {}
pipeline:
node_lines:
- node_line_name: retrieve_node_line
nodes:
- stage: semantic_retrieval
top_k: [4, 8, 16]
modules:
- { component: vectordb, vectordb: faiss_ivf_index, nprobe: [8, 16, 32, 64] }
- node_line_name: post_retrieve_node_line
nodes:
- stage: prompt_maker
modules:
- { component: fstring, prompt: "Question: {query}\n Passage: {retrieved_contents}\n Answer:" }
- stage: generator
modules:
- { component: vllm, model: [Qwen/Qwen2.5-3B-Instruct, Qwen/Qwen2.5-7B-Instruct],
temperature: {range: [0.0, 1.0]} }
See also: configs/rag_stack/config_reference.yaml
(annotated canonical example), and the measured baselines under
configs/baselines/.