Enterprise-grade document intelligence crafted for high-volume pipelines. From scanned bilingual invoices to complex contracts with exact coordinate grounding.
EDUCATION: Indian Institute of Information Technology
B.Tech in Information Technology · CGPA: 9.24
Candidate Name
Keshav Agrawal
Employer
NeoCortex AI
Role
Software Engineering Intern
Enterprise Architecture
A comprehensive, self-hosted document pipeline designed to ingest, validate, and index any file format with absolute transparency.
From initial PDF buffer caching and canvas rendering to Unicode text normalization, zero-shot entity extraction, and pgvector embeddings. Every stage runs idempotently with live console telemetry.
Native Devanagari script recognition with Unicode NFC normalization, Hindi month parsing, and automated VLM vision fallback for distorted scans.
Zero hallucinations. Automated invoice balance verification (Subtotal + Tax = Total), 15-char GSTIN checksum verification, and chronological date checks.
Combine 1536-dimensional semantic vector embeddings with PostgreSQL Full-Text Search and trigram fuzzy matching across your entire document repository.
99.4%
OCR Precision
< 2.5s
Async Processing
100%
Private & Self-Hosted
14 Stages
End-to-End Pipeline
Upload any document to test automated classification, OCR bounding boxes, and instant structured data extraction.
Open Document Workspace