Patent Knowledge Graph & Vector Search Platform
Ingested 2.6M historical patents into PostgreSQL, ArangoDB, and Qdrant, extending searchable prior art back to 1976 and cutting embedding cost 48%.

Situation
The platform's prior-art corpus started at 2002 and its European coverage was incomplete, so analyses silently missed decades of patents that could invalidate a claim.
Task
Extend the searchable corpus back to 1976, close the European coverage gap, and make the whole pipeline reproducible and cost-efficient at multi-million-record scale.
Actions
- Parsed raw USPTO archive files and loaded 2.6M records (1976–2001) into PostgreSQL
- Modeled entities and citations as an ArangoDB knowledge graph
- Chunked and embedded the corpus into a Qdrant vector index for semantic prior-art search
- Migrated a 3.4M-record European corpus to PostgreSQL and put its legal-status refresh on an automated weekly job
- Benchmarked 10 AWS instance configurations before committing to the embedding run
- Shipped every migration with a verification gate and a documented rollback path
Results
2.6M
Patents newly searchable
54.9% → 99.95%
Abstract coverage
4.7M
Legal-status records backfilled
48%
Lower embedding cost
24 days
Embedding runtime (from 60)
What I'd Do Next
Extend the graph with additional citation and family relationships and tune the vector index for higher-recall prior-art retrieval.