Skip to content

Experimental Functions — how a trial is scored

Every evaluated config gets two objective scores — quality (ground-truth metrics on a QA dataset) and performance (latency / throughput) — computed by deliberately decoupled subsystems, so quality can be evaluated once and replayed against different hardware (see Quality store and hardware replay).

Which subsystem computes each objective is selected by three orthogonal YAML switches:

Axis YAML key Values Default
1. Quality backend global.eval_backend static_gt | flashrag static_gt
2. Performance source system.performance_source cost_model | measured cost_model
3. rag_ir mode (CM engine) global.rag_ir_mode static | dynamic dynamic

① Quality backend — global.eval_backend

  • static_gt (default) — the bundled RAG-Stack-Evaluator submodule (rag_stack_evaluator/static_rag_evaluator/). Sequential + ReAct over ONE shared algo_search_space.pipeline.node_lines (no methods blocks): the sequential path runs the node lines (retrieval → rerank/compress → generation) per query; adding the flat dataflow knob pipeline.rag_dataflow lets the optimizer pick the ReAct agentic loop per trial. A branch entry carries that branch's SUB-CONFIG — the declared hierarchy is wired into Ax dependents (nothing knob-specific is hardcoded):
pipeline:
  rag_dataflow:
  - sequential
  - react:
      max_iter: [3, 5]   # REQUIRED on the react branch: max Thought/Action
  #                      # rounds — the loop's only cap
  node_lines: [...]      # ONE shared stage list for both dataflows

The react loop drives ONLY semantic_retrieval / hybrid_retrieval, an optional passage_reranker, and the generator (REACT_ACTIVE_STAGES in rag_search_space.py). A sequential trial deactivates the react sub-knobs, AND a react trial deactivates every sequential-only stage (query_expansion, passage_filter, passage_compressor, prompt_maker) — no search dimensions are spent on dead knobs, and the resolver strips gated-off stages so the cost model never provisions them. Scores with the metric functions under static_rag_evaluator/evaluation/. Zero FlashRAG dependency. - flashrag — FlashRAG-backed evaluator (rag_stack/flashrag_quality_evaluator/, lazy import — needs uv pip install -e FlashRAG). Owns the iterative / agentic methods (algo_search_space.pipeline.methods[].mode, see PIPELINE_MAP in pipeline_factory.py): adaptive_rag, iter_retgen, flare, ircot, self_ask, searchr1, corag, search_o1, react, a_rag. Captures the per-query execution trace so the cost model can replay pipelines whose shape isn't known statically.

global:
  eval_backend: static_gt   # or: flashrag

② Performance source — system.performance_source

  • cost_model (default) — analytical simulation, no GPUs needed at run time. RAGCMAssembly (assembly.py) wraps RAGO's sweep + placement over per-stage cost models (LLM stages via GenZ, FAISS via the calibrated roofline model) and prices each trial through the static or dynamic engine (engines.py) per global.rag_ir_mode (axis ③); the selection policy (system.performance_objectives.selection, e.g. max_throughput) reduces the sweep to one score.
  • measured — deploys real vLLM engines on system.system_design_space.available_gpus and times them (rag_stack_evaluator/static_rag_evaluator/measured/). GPU layout is derived, not searched; every trial tears down all vLLM processes to prevent GPU-memory leaks. Requires global.eval_backend: static_gt, the GPU bounds in system.cm_search_space.{min,max}_num_gpus, and a system.system_design_space block (with available_gpus) — see configs/baselines/ for complete examples.
system:
  performance_source: measured     # or: cost_model (default)
  cm_search_space:
    min_num_gpus: 4
    max_num_gpus: 4
    max_num_gpus_per_stage: 2
  system_design_space:
    batch_size_request: [1, 8, 16, 32]   # one global batch (no per-stage dims)
    vllm_kv_cache_dtype: auto
    available_gpus: [cuda:0, cuda:1, cuda:2, cuda:3]

③ rag_ir mode — global.rag_ir_mode

Only meaningful with performance_source: cost_model: which CM engine prices a trial, and therefore what data it consumes?

  • dynamic (default) — the DYNAMIC engine (RAGCMAssembly.evaluate_dynamic): replays the quality run's recorded quality trace + workflow schema. This pair is the frozen CM input contract (rag_stack/rag_ir/; full reference: rag_cm/README.md) — the CM accepts exactly these two artifacts, with no side channels:
  • The quality trace (canonical envelope: {queries: [{question_id, calls}], provenance?}) — exactly one complete invocation DAG per dataset question, with the ordered per-query component calls (generate / retrieve / rerank / …) and real token counts. Field whitelists are enforced at every layer (unknown keys are errors); both backends produce it. Retired envelopes are not accepted.
  • The WorkflowSchema, whose dataflow_type (sequential | agentic) the engine dispatches on and which carries the corpus statistics (corpus_stats: token stats + per-vectordb index sizes) as a first-class field.

One trace-replay path prices sequential and react/iterative trials on the same scale — this is what makes agentic pipelines cost-modelable at all. - static — the STATIC engine (RAGCMAssembly.evaluate_static): offline TokenStats estimates + config-driven aggregate RAGO inputs. Works without ever running a quality evaluation; sequential dataflow only (an agentic trial has no static shape to price).

global:
  rag_ir_mode: dynamic     # default

Valid combinations

performance_source: cost_model performance_source: measured
eval_backend: static_gt ✅ default; CM engine per rag_ir_mode ✅ measured baselines (configs/baselines/)
eval_backend: flashrag ✅ always trace-replay ❌ rejected by the config validator

All three keys are validated up front (config_validator.py) — unknown values fail fast before any evaluation starts. The config follows a fixed 6-section schema (global / dataset / system / optimizer / eval_backend_setting / algo_search_space) — concepts and syntax rules are documented in config_schema.md, and the annotated reference for every key is configs/rag_stack/config_reference.yaml.