Skip to content

RAG-Stack Documentation

Everything beyond the minimal installation and first-run command in the main README.

  • User guide — the top-level entry point for using RAG-Stack via the rag_stack command: launching a run from a config, monitoring, resuming / recovering an existing project dir, quality stores & hardware replay, system transfer, and the system settings + cache-directory layout every run inherits.
  • Experimental functions — how a trial gets its two objective scores: the three evaluation switches (eval_backend / performance_source / rag_ir_mode), the sequential+react co-mode, the frozen CM input contract (quality trace + workflow schema), and the valid combinations.
  • Config schema — concepts and syntax rules for the 6-section run config (global / dataset / system / optimizer / eval_backend_setting / algo_search_space).
  • Annotated config reference — every config key, annotated inline.
  • Cost model (RAG-CM) — the CM's public entry points (evaluate_dynamic / evaluate_static / evaluate_fixed), the frozen input contract (canonical quality trace + workflow schema + corpus stats), fixed-mode testing, and legacy-trace conversion.
  • Calibration and hardware migration — staged readiness inventory; the CPU-system, workload, GPU-component, runtime, and retrieval-worker profile layers; forbidden E2E/assembly fitting; and the current H100 fixed-point regression guard.
  • FAISS IVF cost model — the roofline-based vector-search cost model (IVF-PQ + IVF-Flat): quick usage, the simulator-benchmark CLI, calibration, and index-string anatomy.
  • Optimizer benchmark — racing the registered optimizers on the historical-surrogate stand-in: the matrix-sweep YAML, output layout, and the reference Pareto cache (pareto-cache / recompute-metrics).
  • Web app — setup, dev servers, and configuration for the apps/ SPA + its FastAPI backend (rag_stack/server/).