AI-Powered Supplier Quotation Analysis, Procurement Decision Support & RAG Knowledge Base
The AI Procurement Assistant is a Streamlit-based web application that uses Large Language Models (OpenAI GPT-4.1-mini), deterministic business rules, and Retrieval-Augmented Generation (RAG) to support procurement professionals throughout the supplier evaluation process.
The application can analyze supplier quotations, extract structured commercial data, identify procurement risks, compare supplier offers, generate negotiation strategies, produce management-ready reports, and build a searchable knowledge base from procurement documents.
The project demonstrates the practical integration of Generative AI into a real-world procurement workflow while combining AI reasoning with explainable business logic and grounded document retrieval.
- Upload TXT and PDF supplier quotations or contracts
- AI-generated document summaries
- Structured procurement data extraction
- Procurement dashboard for extracted information
Automatically evaluates procurement risks using predefined business rules.
Current evaluations include:
- Payment terms
- Incoterms
- Penalty clauses
- Delivery time
- Price increase
- Minimum order quantity
- Critical procurement information
Each analysis generates:
- Risk Score (0–100)
- Risk Level (Low / Medium / High)
- Detailed business rule findings
This combines LLM-based document understanding with deterministic procurement rules so that key risk evaluations remain transparent and reproducible.
Compare multiple supplier quotations simultaneously.
Comparison categories include:
- Price
- Delivery time
- Payment terms
- Incoterms
- Penalty clauses
- Minimum order quantity
- Price validity
- Overall supplier recommendation
Generate AI-assisted supplier negotiation strategies based on:
- Extracted procurement data
- Business rule evaluation
- Risk assessment
Generated outputs include:
- Executive summary
- Negotiation priorities
- Supplier questions
- Counter proposals
- Draft supplier email
Generate management-ready procurement reports containing:
- Executive summary
- Supplier overview
- Commercial terms
- Business rule findings
- Risk assessment
- Procurement concerns
- Negotiation recommendations
Reports can be exported directly as professionally formatted PDF documents.
Upload multiple PDF documents and build a searchable procurement knowledge base using Retrieval-Augmented Generation (RAG).
The RAG pipeline:
- Extracts text from uploaded PDF documents
- Splits documents into overlapping text chunks
- Generates vector embeddings using OpenAI embeddings
- Stores document chunks and metadata in ChromaDB
- Converts user questions into embeddings
- Performs semantic vector search to retrieve the most relevant document chunks
- Provides the retrieved context to GPT-4.1-mini
- Generates answers grounded in the uploaded documents
- Displays the retrieved source documents and chunk references
If the requested information cannot be found in the retrieved document context, the system is instructed to explicitly state that it cannot answer the question based on the provided documents.
PDF Documents
│
▼
Text Extraction
│
▼
Chunking
│
▼
OpenAI Embeddings
│
▼
ChromaDB Vector Store
│
│
├─────────────────────┐
│ │
│ User Question
│ │
│ ▼
│ Question Embedding
│ │
└──────────────► Semantic Search
│
▼
Top-k Chunks
│
▼
GPT-4.1-mini
/ \
▼ ▼
Answer Source References
For learning purposes, the retrieval process was initially implemented manually using OpenAI embeddings, cosine similarity, and top-k retrieval before integrating ChromaDB as the persistent vector store used by the application.
- Python
- Streamlit
- OpenAI API
- GPT-4.1-mini
- OpenAI Embeddings
- ChromaDB
- NumPy
- JSON Structured Output
- ReportLab
- PyPDF
- HTML / CSS
- Custom Business Rule Engine
- Streamlit Session State
ProcurementAssistant/
│
├── app.py
├── rag_utils.py
├── risk_rules.py
├── file_utils.py
├── pdf_utils.py
├── styles.css
├── requirements.txt
├── README.md
├── .gitignore
│
├── prompts/
│ ├── document_analysis.txt
│ ├── document_extraction_json.txt
│ ├── supplier_comparison.txt
│ ├── negotiation_strategy.txt
│ ├── procurement_report.txt
│ └── risk_analysis.txt
│
├── quotation_samples/
│ └── sample supplier documents
│
└── test_rag.py
Local environment variables and the local ChromaDB database are excluded from version control through .gitignore.
Clone the repository:
git clone https://github.com/SaschaK93/ProcurementAssistant.git
cd ProcurementAssistantInstall the required packages:
pip install -r requirements.txtCreate a .env file:
OPENAI_API_KEY=your_api_key
Run the application:
streamlit run app.pySupplier Documents
│
├──────────────────────────────┐
│ │
▼ ▼
Document Analysis RAG Knowledge Base
│ │
▼ ▼
Structured Extraction Chunking
│ │
▼ ▼
Business Rule Engine Embeddings
│ │
▼ ▼
Risk Assessment ChromaDB
│ │
▼ ▼
Supplier Comparison Semantic Retrieval
│ │
▼ ▼
Negotiation Strategy Grounded Q&A
│
▼
Procurement Report
The RAG knowledge base retrieves relevant information from indexed procurement documents and generates grounded answers with retrieved source references.
Implemented functionality:
- TXT and PDF document support
- AI document summarization
- Structured procurement data extraction
- Procurement dashboard
- Deterministic business rule engine
- Supplier quotation comparison
- Negotiation strategy generation
- Procurement report generation
- PDF export
- RAG procurement knowledge base
- Document chunking with overlap
- OpenAI embeddings
- ChromaDB vector storage
- Semantic document retrieval
- Grounded document Q&A
- Retrieved source references
- Unknown-answer guardrail
- Knowledge base rebuild functionality
- Streamlit Cloud deployment
- Error handling
- Session state management
The project was developed incrementally rather than relying on a high-level AI framework from the beginning.
For the RAG implementation, the core retrieval concepts were first implemented manually:
Text
↓
Chunks
↓
Embeddings
↓
Cosine Similarity
↓
Top-k Retrieval
↓
LLM Context
↓
Grounded Answer
After validating the retrieval pipeline, the manual vector search was replaced with ChromaDB for persistent vector storage and semantic retrieval.
This approach provided hands-on understanding of the individual components behind a RAG system before introducing a dedicated vector database.
Current limitations include:
- PDF text extraction does not include OCR for scanned documents
- The RAG knowledge base currently supports PDF documents
- Retrieved source references identify the retrieved document chunks but are not fine-grained inline citations
- Cloud-hosted local vector storage should not be treated as permanent external database storage
- AI-generated recommendations require human review before procurement decisions are made
Planned future development includes:
- Systematic evaluation of AI-assisted vs. manual procurement workflows
- RAG retrieval and answer-quality evaluation
- Agentic procurement workflows
- Tool calling
- LangGraph-based workflow orchestration
- Human-in-the-loop approval steps
- Additional procurement business rules
- OCR support for scanned PDFs
- Multi-language document support
This project is licensed under the MIT License.
Sascha Knies
GitHub: https://github.com/SaschaK93