Open source

More than
side projects.

Built to solve a problem I ran into, or just for the fun of it, then shared as open-source libraries or deployed on the cloud.

01 — Project

codiff

Structural call-graph diff for AI-assisted codebases

★ 7 starsPythonApache-2.0
Python 100%

Install

pip install codiff

Function-level diff

Added, modified and removed functions with sig / calls +N−N annotations, and call chains colour-coded across every file in the diff.

Incremental index

A SQLite call-graph index at the repo root. Only files changed since the last indexed commit are re-parsed, plus the stale callers that pointed at them — so diffs stay fast on large repositories.

Three surfaces

A CLI with terminal, Mermaid and JSON output; a GitHub Action that posts a structural diff on every pull request; and an MCP server for Claude Code, Codex, Gemini CLI and Vibe.

Python and TypeScript

Parses .py, .ts and .tsx, resolving imports and inheritance to build one graph across the whole tree.

Built with

Python 3.11+GitSQLiteMCPGitHub ActionsMermaid

02 — Project

codebot

Semantic code search for coding agents, over a call graph

Python
Python 100%

Install

pip install git+https://github.com/InkyLabAI/codebot_mcp.git

Graph RAG retrieval

Vector search finds entry points; the call graph expands them into a ranked tree of callers and callees, so the agent gets context, not fragments.

Codebase overview

Community detection over the call graph produces a map of the major modules and their entry points — cached after the first call.

Always-fresh index

A file watcher with a 2.5 s debounce re-parses and re-embeds only the functions that changed. Searches block while a re-index runs, so the agent never reads a torn index.

MCP, REST and CLI

The same capabilities through three surfaces, with one-command setup for Claude Code, Cursor, Windsurf, Cline, Copilot and Vibe.

Built with

Python 3.11+Voyage AI embeddingsMCPFastAPICommunity detection

03 — Project

ImageDB

Full-text search over millions of arXiv figures

Python
Python 58%HTML 23%JavaScript 19%

Figure & caption extraction

PyMuPDF worker pool that locates figures and their nearby “Fig. N” captions from text and image layout, with a timeout guard for corrupted PDFs.

Streaming ingestion

Monthly .tar bundles are streamed from arXiv’s requester-pays S3 bucket, prefetching the next bundle while the current one is processed.

Postgres full-text search

Captions indexed with a tsvector + GIN index kept in sync by a trigger, so search stays instant as the table grows.

Cost-aware serving

Images are re-encoded to a high-res JPEG and a low-res WebP on object storage and served through pre-signed URLs — the storage bill, not the database, is the design constraint.

Built with

FastAPIPostgreSQL 17SQLAlchemy (async)PyMuPDFS3 / OVH Object StorageDocker Composenginx

04 — Project

NewsDB

Collect, search and visualise millions of press articles

Python
Python 76%TypeScript 21%Other 3%

Three search modes

Semantic (HNSW over half-precision vectors, result set snapshotted in Redis for stable pagination), keyword (tsvector with keyset pagination) and plain browse — selected automatically from the request.

Local embedding service

Jina v3 multilingual embeddings served through vLLM behind an async FastAPI endpoint, with graceful degradation when no GPU is present.

Built for scale

Keyset pagination, a composite (date, rowid) index, a GIN index for full-text and an HNSW index for vectors — so scrolling stays constant-time at millions of rows.

Clean architecture

Built on the Decorator, Strategy, Registry and Factory patterns, so adding a news site means writing one small class — pagination, deduplication and fetching come for free.

Built with

FastAPICeleryRedisPostgreSQL + pgvectorReact + VitevLLMDocker Compose

There is more on GitHub

Smaller experiments, a WhatsApp booking agent, benchmarks and older work all live on my profile.