Your AI-powered study assistant β chat, quiz, explain, summarize, and more.
PaperBrain turns your documents into an interactive study session. Upload a PDF, ask questions about it, generate a quiz, get flashcards, or request a plain-English explanation β all in one place.
It runs on a FastAPI + React stack with Qwen 2.5-72B via HuggingFace for cloud inference, ChromaDB for document retrieval (RAG), Valkey + BetterDB for semantic caching and observability, and an optional local n8n AI Agent powered by Ollama for fully offline automation.
| Feature | Description | |
|---|---|---|
| π¬ | General Chat | Study assistant backed by Qwen 2.5-72B |
| π | RAG Mode | Ask questions directly about your uploaded documents |
| π§ͺ | Quiz | Auto-generated multiple-choice quizzes on any topic |
| π | Flashcards | Smart cards for active recall and memorization |
| π‘ | Explain | Concept breakdowns at beginner / intermediate / advanced level |
| π | Summarize | Auto-summarize any text or uploaded document |
| π | Document Manager | Upload PDF, TXT, DOCX β indexed per user with isolation |
| π€ | Auth | JWT-based register/login with per-user data separation |
| π | Profile & Stats | Quiz history, streaks, and progression tracking |
| π | n8n AI Agent | Local agent with Ollama (llama3.1 / Qwen 2.5) + 5 tools |
| β‘ | Semantic Cache | Valkey-backed cache that hits on meaning, not exact text β paraphrased questions skip the LLM entirely |
| π | Observability | BetterDB watches Valkey + the cache live; optional Prometheus/Grafana dashboards |
PaperBrain/
βββ backend/ # FastAPI Python backend
β βββ app/
β β βββ auth/
β β β βββ jwt_handler.py # JWT token creation/decoding
β β β βββ middleware.py # get_current_user dependency
β β βββ db/
β β β βββ database.py # SQLite + SQLAlchemy setup
β β β βββ models.py # User, QuizResult, StudySession
β β β βββ crud.py # DB operations
β β βββ tools/ # AI tool modules
β β βββ agent.py # Main AI dispatcher β checks/fills the cache
β β βββ cache.py # Valkey + BetterDB semantic cache (app/cache.py)
β β βββ ingest.py # Document ingestion + chunking
β β βββ rag.py # ChromaDB vector store
β β βββ router_service.py # API routes
β β βββ schemas.py # Pydantic request models
β β βββ main.py # FastAPI app entry point
β βββ Dockerfile
β βββ requirements.txt
βββ frontend/
β βββ src/
β βββ pages/
β βββ Chat.jsx
β βββ Quiz.jsx
β βββ Flashcards.jsx
β βββ Documents.jsx
β βββ Profile.jsx
βββ n8n/ # Local n8n AI Agent
β βββ workflows/
β βββ PaperBrain.json
βββ monitoring/ # Optional Prometheus + Grafana stack
β βββ prometheus.yml # Scrapes BetterDB's /api/prometheus/metrics
β βββ grafana-datasources.yml # Pre-wires the Prometheus datasource
βββ docs/
βββ logo.png
βββ n8n-workflow.png
- Python 3.10+
- Node.js 18+
- Ollama (only required for the n8n local agent)
git clone https://github.com/ApyHtml20/PaperBrain.git
cd PaperBraincd backend
pip install -r requirements.txtCreate a .env file:
SECRET_KEY=your_secret_key_here
# LLM routing (app/llm/) β set at least one; the app cascades through
# whichever of these are present, in this order by default.
GROQ_API_KEY=your_groq_key
HF_TOKEN=your_huggingface_token
OPENAI_API_KEY=your_openai_key
# Optional overrides
GROQ_MODEL=groq/llama-3.3-70b-versatile
HF_MODEL=Qwen/Qwen2.5-72B-Instruct
OPENAI_MODEL=gpt-4o-mini
LLM_FALLBACK_ORDER=groq,huggingface,openai
# Cache (app/cache.py) β needs a running Valkey, see step 4 below.
# Safe to leave everything at its default: a Valkey outage just disables the
# cache and falls back to normal (uncached) LLM calls.
VALKEY_HOST=localhost
VALKEY_PORT=6379
SEMANTIC_CACHE_ENABLED=true
SEMANTIC_CACHE_THRESHOLD=0.12
SEMANTIC_CACHE_TTL=86400Start Valkey (needed even outside Docker Compose β see step 4):
docker run -d -p 6379:6379 valkey/valkey:8-alpineStart the server:
uvicorn app.main:app --reload --port 8000cd frontend
npm install
npm run dev# Install and start Ollama, then pull the model
ollama pull llama3.1
# Install and start n8n
npm install -g n8n
n8n start
# β http://localhost:5678Then in n8n:
- Go to Workflows β Import
- Import
n8n/workflows/PaperBrain.json - Set the Ollama node URL to
http://localhost:11434 - Click Publish
The whole stack β backend, frontend, Valkey, and BetterDB Monitor β runs with one command:
docker compose up -d| Service | URL | What it's for |
|---|---|---|
| BetterDB Monitor | http://localhost:3001 | Live Valkey health (memory, slowlog, hot keys) + semantic-cache hit-rate, cost saved, and threshold recommendations |
| Backend API | http://localhost:8000 | FastAPI |
| Frontend | http://localhost:5173 | React app |
Add --profile monitoring to also stand up Prometheus + Grafana, scraping BetterDB's betterdb_* and semantic_cache_* metrics into persistent dashboards:
docker compose --profile monitoring up -d| Service | URL | Default login |
|---|---|---|
| Grafana | http://localhost:3002 | admin / admin (override with GRAFANA_PASSWORD) |
| Prometheus | http://localhost:9090 | β |
Running the backend outside Docker (step 2 above)? Point it at any Valkey instance with VALKEY_HOST / VALKEY_PORT β see the .env block in step 2.
The local n8n agent orchestrates all learning tools automatically using Ollama llama3.1 β no cloud required.
[Postman / Frontend]
β
[AI Agent (RAG)] ββ [Ollama β llama3.1]
β
βββββββ΄ββββββββββββββββββββββββββββββββββ
β β β β β
[Flashcards] [Explain] [RAG] [Summarize] [Quiz]
β
[Frontend Output]
| Tool | What it does |
|---|---|
| π Flashcards | Generate flashcards on any topic |
| π‘ Explain | Explain a concept at any depth |
| π RAG | Search your uploaded documents |
| π Summarize | Summarize topics automatically |
| π§ͺ Quiz | Generate multiple-choice questions |
Five of the six AI actions (rag-qa, quiz, flashcards, explain, resume) go through a semantic cache before they go anywhere near an LLM. Unlike a plain key-value cache, it hits on meaning: "C'est quoi la mitose ?" and "Peux-tu m'expliquer la mitose ?" embed to nearly the same vector, so the second question is served from Valkey in milliseconds instead of triggering another Groq/HuggingFace/OpenAI call. chat is deliberately excluded β its answer depends on recent conversation history that a single cached query can't represent, so caching it could replay a reply from an unrelated conversation.
flowchart LR
U([User]) --> FE[React frontend]
FE --> API["FastAPI /api/*"]
API --> AG["agent.py β run_agent()"]
AG -->|1 - embed + check| SC["app/cache.py\nSemanticCache"]
SC -->|hit| AG
AG -.->|cache hit: skip LLM| RESP1([Response])
AG -->|2 - miss| RAG["rag.py β hybrid BM25\n+ embeddings retrieval"]
RAG --> LLM["llm/router.py\nLiteLLM: Groq β HF β OpenAI"]
LLM --> AG
AG -->|3 - store result| SC
AG --> RESP2([Response])
SC <==> VK[(Valkey)]
RAG -.-> CHROMA[(ChromaDB)]
BDB["BetterDB Monitor\n:3001"] -->|watches| VK
BDB -->|/api/prometheus/metrics| PROM["Prometheus\n:9090 (optional)"]
PROM --> GRAF["Grafana\n:3002 (optional)"]
style SC fill:#4f46e5,color:#fff
style VK fill:#dc2626,color:#fff
style BDB fill:#059669,color:#fff
Cache hit vs. cache miss:
sequenceDiagram
participant U as User
participant A as agent.py
participant C as Valkey / SemanticCache
participant L as LLM router
U->>A: "Explique la mitose"
A->>C: embed + check(prompt)
alt cache hit (similar question seen before)
C-->>A: cached result (cosine distance < threshold)
A-->>U: instant response, cached: true
else cache miss
C-->>A: no match
A->>L: complete(system, prompt)
L-->>A: LLM answer
A->>C: store(prompt, result)
A-->>U: response, cached: false
end
- Valkey β Linux Foundation fork of Redis, protocol-compatible, open-source license. Drop-in cache backend, nothing app-specific about it.
- BetterDB β purpose-built to observe exactly this pairing: it watches Valkey directly (slowlog, hot keys, memory, ops/sec) and auto-discovers every
SemanticCacheinstance registered in Valkey's__betterdb:cacheshash, so hit-rate, cost saved, and similarity-threshold recommendations show up in its dashboard with zero extra wiring.
rag-qa can echo a user's own uploaded documents back in the answer, so each user gets their own cache namespace β same isolation model as the per-user ChromaDB collections. quiz, flashcards, explain, and resume only ever depend on a topic string, nothing user-specific, so those are cached globally: the first person who asks for a "quiz on photosynthesis" pays the LLM cost, everyone after gets it free. See SHARED_CACHE_ACTIONS in backend/app/cache.py.
| Variable | Default | Purpose |
|---|---|---|
VALKEY_HOST / VALKEY_PORT |
valkey / 6379 (Docker) β localhost / 6379 (local) |
Where the backend finds Valkey |
SEMANTIC_CACHE_ENABLED |
true |
Kill switch β false bypasses the cache entirely |
SEMANTIC_CACHE_THRESHOLD |
0.12 |
Max cosine distance for a hit; lower = stricter matching |
SEMANTIC_CACHE_TTL |
86400 (24h) |
How long a cached answer stays valid |
BETTERDB_ENCRYPTION_KEY |
dev placeholder | At-rest encryption key for BetterDB's stored credentials β set a real one in production |
GRAFANA_PASSWORD |
admin |
Grafana admin password (monitoring profile only) |
A Valkey outage never breaks the app β app/cache.py swallows connection errors and falls back to an uncached LLM call, so caching is strictly an optimization, never a hard dependency.
| Method | Route | Description |
|---|---|---|
POST |
/api/auth/register |
Create a new account |
POST |
/api/auth/login |
Log in and receive a JWT token |
| Method | Route | Description |
|---|---|---|
POST |
/api/chat |
General AI chat |
POST |
/api/rag-qa |
Chat with your documents |
POST |
/api/quiz |
Generate an MCQ quiz |
POST |
/api/flashcards |
Generate flashcards |
POST |
/api/explain |
Explain a concept |
POST |
/api/resume |
Summarize text |
| Method | Route | Description |
|---|---|---|
GET |
/api/documents |
List your documents |
POST |
/api/upload |
Upload a PDF, TXT, or DOCX file |
DELETE |
/api/documents/{filename} |
Delete a document |
FROM python:3.11-slim
WORKDIR /code
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "7860"]Add these under Settings β Variables and secrets:
HF_TOKEN=...
HF_MODEL=Qwen/Qwen2.5-72B-Instruct
SECRET_KEY=...
Backend β FastAPI Β· SQLite + SQLAlchemy Β· ChromaDB (hybrid BM25 + multilingual embeddings, per-user collections) Β· LiteLLM (Groq / HuggingFace / OpenAI cascade) Β· Valkey + BetterDB semantic cache Β· python-jose Β· pdfplumber Β· python-docx
Observability β BetterDB Monitor (Valkey health + cache analytics) Β· Prometheus Β· Grafana (optional, monitoring profile)
Frontend Gradio
Local AI β n8n Β· Ollama Β· llama3.1
Deployment β Hugging Face Spaces (Docker)
- Passwords hashed with SHA-256 + random salt
- JWT tokens expire after 24 hours
- All routes protected by auth middleware
- Documents and ChromaDB collections isolated by
user_id - Files stored under
documents/{user_id}/


