Skip to content

Repository files navigation

TendorBot Pakistan 🇵🇰

AI-Powered PPRA Tender Automation Agent

TendorBot is an AI-powered tender intelligence and compliance assistant designed to help Pakistani businesses analyze PPRA tenders, extract requirements, verify eligibility, identify missing documents, and automate tender preparation.

The system combines Retrieval-Augmented Generation (RAG), semantic similarity search, document processing, and LLM-powered workflow automation to turn lengthy tender documents into actionable compliance insights.

Core Workflow

Tender PDF → Requirement Extraction → Semantic Matching → Eligibility Analysis → Document Generation

Key Features

  • 📄 PPRA tender PDF parsing and analysis
  • 🔎 RAG-based retrieval from tender documents
  • 🧠 LLM-powered requirement extraction
  • 📋 Structured tender_requirements.json intermediate artifact
  • 🏢 Company document knowledge base
  • 📊 Semantic matching between tender requirements and company credentials
  • ⚠️ Automated eligibility and compliance-gap detection
  • 📑 Missing-document identification
  • ✍️ AI-generated cover letters and submission checklists
  • 🌐 English + Urdu support
  • ⚡ Fast LLM inference with Groq

Technology Stack

  • Frontend: Streamlit
  • PDF Processing: pypdf
  • Document Processing: pypdf, python-docx, and text processing
  • Vector Database: FAISS and ChromaDB -- FAISS for tender retrieval, ChromaDB for company document knowledge base
  • Embeddings: Sentence Transformers all-MiniLM-L6-v2 and Google Gemini gemini-embedding-001
  • RAG: Retrieval-Augmented Generation for tender and company document knowledge bases
  • LLM: Groq / Llama 3.3 70B and Google Gemini
  • Semantic Matching: Embedding-based similarity matching between tender requirements and company credentials
  • Programming Language: Python
  • Languages Supported: English + Urdu

Company Knowledge Base

It keeps track of company documents (NTN, PEC, audit reports, experience letters) and answers questions using only information from those documents. If the required information is not found in the documents, it does not call the LLM, completely preventing hallucinations!

Notes

  • One Chroma collection per company_id (data isolation)
  • Re-uploading the same file is safe (content-hash ids -> upsert)
  • Scanned/image PDFs have no text layer -> clear error message (OCR not supported)
  • Thresholds MIN_SIMILARITY=0.55 / strong 0.75 are tunable via env
  • Uses the current google-genai SDK (the older google-generativeai package is deprecated)

About

AI-powered PPRA tender analysis & compliance automation using RAG and semantic matching.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages