This repository (LoCoRAG) is the implementation of Self-Correcting Agentic RAG via Memory-Grounded Failure Localization.
Create the Python environment:
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtIf your machine needs GPU FAISS, install the FAISS package that matches your CUDA environment instead of faiss-cpu.
Copy the environment template:
cp .env.example .envConfigure the LLM provider in .env.
For OpenAI-compatible APIs:
AI_PROVIDER=openai
OPENAI_API_KEY=your_api_key
OPENAI_BASE_URL=https://api.openai.com/v1
OPENAI_LLM_MODEL=gpt-4.1-miniFor vLLM:
AI_PROVIDER=vllm
VLLM_HOST=http://127.0.0.1
VLLM_LLM_PORT=30023
VLLM_LLM_MODEL=your_served_model_nameThe RAG pipeline expects a retriever service at:
SEARCH_SERVICE_HOST=127.0.0.1
SEARCH_SERVICE_PORT=8091
SEARCH_SERVICE_ENDPOINT=/retrieve
TOP_K=5
FACT_EXTRACTION_MODE=two_stageThe included local retriever uses an E5 encoder plus a FAISS index. Small corpus files are provided in retriever/Corpus/.
Build an index:
CORPUS_FILE=retriever/Corpus/hotpotqa_corpus.jsonl \
INDEX_SAVE_DIR=retriever/indexes/hotpotqa \
RETRIEVER_MODEL=intfloat/e5-large-v2 \
bash retriever/build_index.shStart the retriever:
INDEX_FILE=retriever/indexes/hotpotqa/e5-large_Flat.index \
CORPUS_FILE=retriever/Corpus/hotpotqa_corpus.jsonl \
RETRIEVER_MODEL=intfloat/e5-large-v2 \
bash retriever/retrieval.shThe retriever serves POST /retrieve on port 8091.
Use the matching dataset and corpus/index pair for other tasks, for example:
CORPUS_FILE=retriever/Corpus/musique_corpus.jsonl \
INDEX_SAVE_DIR=retriever/indexes/musique \
bash retriever/build_index.shThen start it with:
INDEX_FILE=retriever/indexes/musique/e5-large_Flat.index \
CORPUS_FILE=retriever/Corpus/musique_corpus.jsonl \
bash retriever/retrieval.shMake sure the LLM service and retriever are both running, then test one query:
python -c "from rag_pipeline_lib import llm_adapter; from rag_pipeline_lib.pipeline import run_multistep_pipeline; llm_adapter.configure_llm_provider(); print(run_multistep_pipeline('When did Lothair II mother die?', verbose=True))"Input JSONL format:
{"id": "case_id", "input": "question", "output": [{"answer": "gold answer"}]}Run batch prediction:
python evaluation/generate_predictions_from_multistep.py \
data/hotpotqa.jsonl \
Result/hotpotqa_predictions.jsonl \
--sample_size 50 \
--sequential_sampling \
--max_workers 4Resume an interrupted run:
python evaluation/generate_predictions_from_multistep.py \
data/hotpotqa.jsonl \
Result/hotpotqa_predictions.jsonl \
--resume \
--max_workers 4Prediction output format:
{"id": "case_id", "question": "question", "output": [{"answer": "predicted answer", "reasoning": "optional reasoning"}]}Compute exact match, F1, cover EM, and ROUGE-L:
python evaluation/evaluation_script.py \
data/hotpotqa.jsonl \
--guess_file Result/hotpotqa_predictions.jsonlWrite per-example metrics:
python evaluation/evaluation_with_metrics.py \
data/hotpotqa.jsonl \
--guess_file Result/hotpotqa_predictions.jsonl \
--output_file Result/hotpotqa_predictions_with_metrics.jsonlRun LLM-based answer judging:
python evaluation/llm_evaluate_predictions.py \
--prediction_file Result/hotpotqa_predictions.jsonl \
--dataset_file data/hotpotqa.jsonl \
--output_file Result/hotpotqa_llm_eval.jsonl \
--max_workers 8| File | Description |
|---|---|
data/hotpotqa.jsonl |
HotpotQA-style QA subset |
data/2wikimultihopqa.jsonl |
2WikiMultihopQA-style QA subset |
data/musique.jsonl |
MuSiQue-style QA subset |
data/ragtracer.jsonl |
RAG tracing / QA subset |
retriever/Corpus/hotpotqa_corpus.jsonl |
HotpotQA retriever corpus subset |
retriever/Corpus/2wikimultihopqa_corpus.jsonl |
2Wiki retriever corpus subset |
retriever/Corpus/musique_corpus.jsonl |
MuSiQue retriever corpus subset |
Runtime outputs are ignored by Git, including Result/, data/evidence_memory/, data/log.txt, and retriever/indexes/.