Experimental Functions — how a trial is scored¶
Every evaluated config gets two objective scores — quality (ground-truth metrics on a QA dataset) and performance (latency / throughput) — computed by deliberately decoupled subsystems, so quality can be evaluated once and replayed against different hardware (see Quality store and hardware replay).
Which subsystem computes each objective is selected by three orthogonal YAML switches:
| Axis | YAML key | Values | Default |
|---|---|---|---|
| 1. Quality backend | global.eval_backend |
static_gt | flashrag |
static_gt |
| 2. Performance source | system.performance_source |
cost_model | measured |
cost_model |
| 3. rag_ir mode (CM engine) | global.rag_ir_mode |
static | dynamic |
dynamic |
① Quality backend — global.eval_backend¶
static_gt(default) — the bundled RAG-Stack-Evaluator submodule (rag_stack_evaluator/static_rag_evaluator/). Sequential + ReAct over ONE sharedalgo_search_space.pipeline.node_lines(nomethodsblocks): the sequential path runs the node lines (retrieval → rerank/compress → generation) per query; adding the flat dataflow knobpipeline.rag_dataflowlets the optimizer pick the ReAct agentic loop per trial. A branch entry carries that branch's SUB-CONFIG — the declared hierarchy is wired into Axdependents(nothing knob-specific is hardcoded):
pipeline:
rag_dataflow:
- sequential
- react:
max_iter: [3, 5] # REQUIRED on the react branch: max Thought/Action
# # rounds — the loop's only cap
node_lines: [...] # ONE shared stage list for both dataflows
The react loop drives ONLY semantic_retrieval / hybrid_retrieval, an
optional passage_reranker, and the generator
(REACT_ACTIVE_STAGES in
rag_search_space.py). A
sequential trial deactivates the react sub-knobs, AND a react trial
deactivates every sequential-only stage (query_expansion,
passage_filter, passage_compressor, prompt_maker) — no search
dimensions are spent on dead knobs, and the resolver strips gated-off
stages so the cost model never provisions them. Scores with the metric
functions under
static_rag_evaluator/evaluation/.
Zero FlashRAG dependency.
- flashrag — FlashRAG-backed evaluator
(rag_stack/flashrag_quality_evaluator/,
lazy import — needs uv pip install -e FlashRAG). Owns the iterative /
agentic methods (algo_search_space.pipeline.methods[].mode, see
PIPELINE_MAP in
pipeline_factory.py):
adaptive_rag, iter_retgen, flare, ircot, self_ask, searchr1,
corag, search_o1, react, a_rag. Captures the per-query execution
trace so the cost model can replay pipelines whose shape isn't known
statically.
② Performance source — system.performance_source¶
cost_model(default) — analytical simulation, no GPUs needed at run time.RAGCMAssembly(assembly.py) wraps RAGO's sweep + placement over per-stage cost models (LLM stages via GenZ, FAISS via the calibrated roofline model) and prices each trial through the static or dynamic engine (engines.py) perglobal.rag_ir_mode(axis ③); the selection policy (system.performance_objectives.selection, e.g.max_throughput) reduces the sweep to one score.measured— deploys real vLLM engines onsystem.system_design_space.available_gpusand times them (rag_stack_evaluator/static_rag_evaluator/measured/). GPU layout is derived, not searched; every trial tears down all vLLM processes to prevent GPU-memory leaks. Requiresglobal.eval_backend: static_gt, the GPU bounds insystem.cm_search_space.{min,max}_num_gpus, and asystem.system_design_spaceblock (withavailable_gpus) — see configs/baselines/ for complete examples.
system:
performance_source: measured # or: cost_model (default)
cm_search_space:
min_num_gpus: 4
max_num_gpus: 4
max_num_gpus_per_stage: 2
system_design_space:
batch_size_request: [1, 8, 16, 32] # one global batch (no per-stage dims)
vllm_kv_cache_dtype: auto
available_gpus: [cuda:0, cuda:1, cuda:2, cuda:3]
③ rag_ir mode — global.rag_ir_mode¶
Only meaningful with performance_source: cost_model: which CM engine
prices a trial, and therefore what data it consumes?
dynamic(default) — the DYNAMIC engine (RAGCMAssembly.evaluate_dynamic): replays the quality run's recorded quality trace + workflow schema. This pair is the frozen CM input contract (rag_stack/rag_ir/; full reference: rag_cm/README.md) — the CM accepts exactly these two artifacts, with no side channels:- The quality trace (canonical envelope:
{queries: [{question_id, calls}], provenance?}) — exactly one complete invocation DAG per dataset question, with the ordered per-query component calls (generate / retrieve / rerank / …) and real token counts. Field whitelists are enforced at every layer (unknown keys are errors); both backends produce it. Retired envelopes are not accepted. - The
WorkflowSchema, whosedataflow_type(sequential|agentic) the engine dispatches on and which carries the corpus statistics (corpus_stats: token stats + per-vectordb index sizes) as a first-class field.
One trace-replay path prices sequential and react/iterative trials on the
same scale — this is what makes agentic pipelines cost-modelable at all.
- static — the STATIC engine (RAGCMAssembly.evaluate_static): offline
TokenStats estimates + config-driven aggregate RAGO inputs. Works without
ever running a quality evaluation; sequential dataflow only (an agentic
trial has no static shape to price).
Valid combinations¶
performance_source: cost_model |
performance_source: measured |
|
|---|---|---|
eval_backend: static_gt |
✅ default; CM engine per rag_ir_mode |
✅ measured baselines (configs/baselines/) |
eval_backend: flashrag |
✅ always trace-replay | ❌ rejected by the config validator |
All three keys are validated up front
(config_validator.py) — unknown values fail
fast before any evaluation starts. The config follows a fixed 6-section
schema (global / dataset / system / optimizer /
eval_backend_setting / algo_search_space) — concepts and syntax rules are
documented in config_schema.md, and the annotated
reference for every key is
configs/rag_stack/config_reference.yaml.