Not just prompting. I build AI systems that go beyond chatbots.
Not wrappers. Not demos. Real architectures: multi-agent coordination, novel RAG retrieval, voice-vision pipelines, fine-tuned models, autonomous build pipelines shipped end-to-end with tests, CI, and evals.
Building in public. Looking for AI engineer / applied AI internship roles.
Agent Factory - 6-Agent Pipeline That Ships Full Projects
One command,
/forge, and a team of six specialized agents (idea-hunter, architect, backend-engineer, frontend-engineer, reviewer, debugger) turns nothing into a tested, runnable project.
Not a code generator. A pipeline with contracts. The architect freezes an API contract before backend and frontend build in parallel from it - that's what keeps two agents' output compatible without either seeing the other's code. The reviewer is read-only by design; the debugger is the only agent that runs code, graded on fixing what it finds rather than writing more.
Agents: idea-hunter -> architect -> backend + frontend (parallel) -> reviewer -> debugger -> devops-engineer
Proof: 3 full projects shipped end-to-end in test runs - PinPoint, DriftGuard, Receipts.dev
Receipts.dev: 92 files, 16 bugs found and fixed (1 critical, 4 high), 0 type errors, 0 lint errors
Contract: runs/<timestamp>/ file handoff - architect's API contract is the single source of truth
Roadmap: Phase 1 (Claude Code subagents, live) -> Phase 2 (Python SDK orchestrator, factory.py, live)
Claude Code Subagents Anthropic Python SDK Multi-Agent Orchestration Streaming Tool Loop ThreadPoolExecutor
Receipts.dev - Prove Skills With Code, Not Buzzwords
AI-powered skill verification from real Git history. Every skill on the profile deep-links to the actual commit that proves it, via a recruiter chat that can only cite real diffs and never invent a claim.
Built by Agent Factory's full pipeline in a single run: idea to architecture to parallel backend/frontend to review to debug. GitHub OAuth with Fernet-encrypted tokens, an async GitHub client with retry, a pgvector code-chunk retriever, and grounded chat with hallucination-proof citation validation.
Next.js 15 FastAPI pgvector GitHub OAuth Celery Grounded RAG
Personal LLM - Local-First Memory + RAG Kernel
One memory engine, built once, imported by everything else: a local-first, privacy-preserving memory + RAG core (SQLite + ChromaDB + hybrid model router) that answers with citations and refuses honestly when it doesn't know.
Not a demo - infrastructure. Three downstream apps import it instead of rebuilding retrieval: second-brain (vault ingestion, auto-linking, offline knowledge-graph viewer - 40 tests) and github-pr-agent (repo analysis, issue triage, PR planning - 32 tests) are public; DreamOS (an Electron AI command bar over the same engine) ships when its demo video does.
Tests: 100 offline, fully mocked - zero-key CI on every push
Agent layer: plan-act-reflect loop, 4 permission-tiered tools (incl. SSRF-guarded web fetch), full audit log
Voice/vision: faster-whisper STT (offline, free) + OCR ingestion; native-crash inputs pre-validated (PyAV)
Security: token-authenticated HTTP gateway; browser-Origin requests rejected outright
Python FastAPI ChromaDB SQLite sentence-transformers Ollama Gemini faster-whisper
CivilizationOS - Multi-Agent AI Society
A living simulation: 10 autonomous citizen-agents + 5 institutional councils (35 AI agents) debate, remember, and react to injected crises - Pandemic, Drought, Cyberattack, Election, Crime Wave, and now self-generated emergent crises.
Novel contribution - Temporal-Causal Memory Fusion (TCMF): Standard RAG retrieves by semantic similarity alone. TCMF fuses two streams:
AGORA stream - citizen episodic memories scored by relevance x recency x importance
PANTHEON stream - societal causal graph (NetworkX DiGraph): crisis -> decision -> outcome
Fused score = episodic_score(m, q) x (1 + lambda x causal_boost(m))
A witness to a root cause outranks someone who heard about it second-hand. No off-the-shelf RAG system does this. Full design write-up with code and tradeoffs: docs/tcmf.md
Validated it with a controlled benchmark against 6 baseline retrieval strategies: the original multiplicative fusion scored 0.00 on the causal signal it was meant to exploit, a fixed normalized-additive version recovered it and beat every single-signal baseline in a mixed regime. A follow-up 60-scenario study showed retrieval choice changes the model's actual decision, not just its confidence.
Latest additions: sustained-fear auto-crisis injection so the society generates its own emergencies, per-council effectiveness scoring (debate to verdict to 60-tick fear delta), union-find citizen faction detection on mutual affinity, and a Story Rewind scrubber over the full causal timeline.
3-tier LLM router: Ollama/Qwen2.5-3B ($0) -> Gemini Flash ($0) -> Claude API (~$0.002/debate)
Fine-tuning: LoRA on Qwen2.5-3B | MLflow tracking | persona-consistency eval harness
Full-stack: FastAPI + WebSocket <-> React + Three.js 3D city (replaced the earlier PixiJS UI)
Tests: 61 passing
Total cost: Under $5 to build.
Live now: civilization-os-murex.vercel.app (frontend) - backend on Render.
Python TypeScript FastAPI React Three.js ChromaDB Ollama Gemini Claude LoRA MLflow NetworkX
Recall - Spatial AI Memory
Point your phone camera at your space. Ask out loud "where did I leave my keys?" Get a spoken answer with the exact frame it was seen in.
Not another AI wrapper. Persistent spatial memory across sessions. Time-decay re-ranking. Gemini Live function-calling into local ChromaDB. The voice model doesn't hallucinate locations - it calls a tool that searches a vector store built from what the camera actually saw.
Eval (June 2026): Recall@1 100% (10/10) | Recall@3 100% (10/10) | Median latency 149ms
Embeddings: all-MiniLM-L6-v2 via ONNX - fully local, zero embedding cost
Voice: Gemini Live push-to-talk with function calling
Quota management: 120s minimum between vision calls + daily budget counter on-screen
Total commits: 162
Live now: deployed on Render.
Python FastAPI React ChromaDB Gemini Live ONNX WebSocket cloudflared
resume-job-fit-ai - AI Resume Scorer | Live Demo
Fit scoring, keyword analysis, AI-rewritten bullet diffs, multi-tone cover letter, interview prep, skills gap roadmap, LinkedIn optimizer - one click.
Deployed: Streamlit Community Cloud (live now, free tier, no credit card)
Tests: 29 unit tests | GitHub Actions CI on every push
Outputs: Pydantic-validated structured JSON - no brittle string parsing
Features: 12+ tools: multi-job comparison, application tracker, cover letter, DOCX export
Python Gemini Streamlit Pydantic SQLite pdfplumber GitHub Actions
Two small studies, each with a negative control and a replication on a second model family, because a finding from one model is a fact about that model.
AUGUR - Where a time-locked model actually stands
A model trained on nothing published after 1938, asked what year it is, answers 1899.
Vintage language models are being used to reconstruct historical belief, on the assumption that a 1938 cutoff represents 1938. It does not. A cutoff describes the edge of a corpus, not its centre of gravity, and both models tested answer from four to eight decades earlier than their stated cutoff. Telling one of them the year relocates it forty years on the spot; telling the other does nothing at all - so a time-locked model has to be characterised on two axes, where it stands and whether it can be moved.
A second finding fell out of the demo: asked whether another great war was coming, a 1930-anchored model says no and calls it highly improbable. Asked instead to enumerate the dangers, the same model in the same year gives an accurate threat assessment naming the Rhineland, Poland, and Czechoslovakia. Both were in the corpus. The form of the question decides which 1930 you meet.
n=2 replication two model families 14-probe battery, temperature 0, fixed seed
COLD READ - The anonymity half-life belongs to the reader
How many words can you write before a machine knows who you are? The question is malformed, and why it is malformed is the finding.
72 authors from a labelled blog corpus, shown to local models in growing slices, guessing gender, age band, and star sign at every step. Star sign is the negative control - labelled in the data, not inferable from prose - and it never left the floor in either model, which is what makes the rest trustworthy.
Same authors, same words, same prompt, two readers: one model needs the better part of a thousand words to beat a coin flip on gender, the other needs fifty. A sixteen-fold gap on identical text. No statement of the form "you are anonymous for N words" means anything without naming the model. Corpus memorisation tested directly and ruled out. Ships with a two-seat consent-gated app that profiles both participants live.
n=2 replication negative control contamination ruled out Wilson intervals
Zabira Academy case study - 27 merged PRs into a live product
Twelve days, +8,838/-1,619 across 557 file changes, into a production learning platform I did not build, every change reviewed and merged by the repository owner. Security hardening (exception leakage across 226 files, Origin checks on 60 mutating endpoints, a sanitizer bypass, session revocation), an end-to-end suite covering the seven launch-critical flows with a CI gate, and a sitewide search feature built from backend to voice input.
Including the one where CI went red on my feature branch and the cause turned out to predate my work by three days - a registration validation rule that silently failed a test fixture, bisected through the run history and confirmed from the failed run's Playwright snapshot.
The codebase is private, so the case study is the record.
I run a nightly agent across my own 12-repo project fleet: it triages the next task, implements it, runs the tests, and opens a PR - every change gated by CI and a human merge, nothing lands unreviewed. Shipped and merged 30+ PRs in the past month. Coordination records (task queue, run logs) live in a private repo since they're internal tooling, not a product.
jd/tenacity #668 - MERGED. Fixed static typing for @retry-decorated instance methods so bound-method signatures survive the decorator under mypy strict.
dry-python/returns #2480 - MERGED. Relaxed future() / future_safe() parameter types from Coroutine to the broader Awaitable protocol.
memgraph/gqlalchemy #390 - open PR adding unary-operator support (IS NOT NULL) to the query builder: 4 tests, docs, CI green, CLA signed. Found and scoped with my own triage pipeline (AutoCTO), implemented keylessly via git + the GitHub CLI.
memgraph/gqlalchemy #392 - open PR adding ON CREATE / ON MATCH clauses to the same query builder, with tests and docs.
jazzband/django-taggit #944 - open PR adding remove_by_slug() to TaggableManager, with tests.
google/adk-python #6190 - fixed an Optional[List[str]] type hint bug in cleanup_unused_files that broke the CLI parser (labeled "good first issue" by Google's ADK team). Survived a full maintainer review round - root-caused a CI failure to a leftover repro script breaking Mypy and the pyink linter, fixed it, and got an LGTM - before a maintainer's own commit resolved the same underlying bug first, closing the PR unmerged.
ai_ml = ["RAG architectures", "multi-agent systems", "LoRA fine-tuning",
"vector DBs", "LLM orchestration", "structured outputs", "evals",
"agent pipelines with frozen API contracts"]
apis = ["Gemini", "Claude (Anthropic)", "Ollama", "Gemini Live"]
backend = ["Python 3.11+", "FastAPI", "WebSocket", "Node.js"]
frontend = ["React", "TypeScript", "Vite", "Three.js", "PixiJS"]
infra = ["AWS", "Google Cloud", "Docker", "Streamlit Cloud", "cloudflared", "Vercel"]
tracking = ["MLflow", "Pydantic", "ChromaDB", "SQLite", "GitHub Actions CI"]- Portfolio - zaidalisyed.vercel.app | source (Next.js 16, Three.js WebGL, GSAP)
- Building in public - LinkedIn
- Open to AI engineer internships, applied AI roles, early-stage startups
- Next: Agent Factory Phase 2 (Python SDK orchestrator) hardening, open-sourcing CivilizationOS fully


