Open source
More than
side projects.
Built to solve a problem I ran into, or just for the fun of it, then shared as open-source libraries or deployed on the cloud.
01 — Project
codiff
Structural call-graph diff for AI-assisted codebases
Install
pip install codiff
Function-level diff
Added, modified and removed functions with sig / calls +N−N annotations, and call chains colour-coded across every file in the diff.
Incremental index
A SQLite call-graph index at the repo root. Only files changed since the last indexed commit are re-parsed, plus the stale callers that pointed at them — so diffs stay fast on large repositories.
Three surfaces
A CLI with terminal, Mermaid and JSON output; a GitHub Action that posts a structural diff on every pull request; and an MCP server for Claude Code, Codex, Gemini CLI and Vibe.
Python and TypeScript
Parses .py, .ts and .tsx, resolving imports and inheritance to build one graph across the whole tree.
Built with
02 — Project
codebot
Semantic code search for coding agents, over a call graph
Install
pip install git+https://github.com/InkyLabAI/codebot_mcp.git
Graph RAG retrieval
Vector search finds entry points; the call graph expands them into a ranked tree of callers and callees, so the agent gets context, not fragments.
Codebase overview
Community detection over the call graph produces a map of the major modules and their entry points — cached after the first call.
Always-fresh index
A file watcher with a 2.5 s debounce re-parses and re-embeds only the functions that changed. Searches block while a re-index runs, so the agent never reads a torn index.
MCP, REST and CLI
The same capabilities through three surfaces, with one-command setup for Claude Code, Cursor, Windsurf, Cline, Copilot and Vibe.
Built with
03 — Project
ImageDB
Full-text search over millions of arXiv figures
Figure & caption extraction
PyMuPDF worker pool that locates figures and their nearby “Fig. N” captions from text and image layout, with a timeout guard for corrupted PDFs.
Streaming ingestion
Monthly .tar bundles are streamed from arXiv’s requester-pays S3 bucket, prefetching the next bundle while the current one is processed.
Postgres full-text search
Captions indexed with a tsvector + GIN index kept in sync by a trigger, so search stays instant as the table grows.
Cost-aware serving
Images are re-encoded to a high-res JPEG and a low-res WebP on object storage and served through pre-signed URLs — the storage bill, not the database, is the design constraint.
Built with
04 — Project
NewsDB
Collect, search and visualise millions of press articles
Three search modes
Semantic (HNSW over half-precision vectors, result set snapshotted in Redis for stable pagination), keyword (tsvector with keyset pagination) and plain browse — selected automatically from the request.
Local embedding service
Jina v3 multilingual embeddings served through vLLM behind an async FastAPI endpoint, with graceful degradation when no GPU is present.
Built for scale
Keyset pagination, a composite (date, rowid) index, a GIN index for full-text and an HNSW index for vectors — so scrolling stays constant-time at millions of rows.
Clean architecture
Built on the Decorator, Strategy, Registry and Factory patterns, so adding a news site means writing one small class — pagination, deduplication and fetching come for free.
Built with
There is more on GitHub
Smaller experiments, a WhatsApp booking agent, benchmarks and older work all live on my profile.