Skip to primary content
Technical Reference Architecture

Autonomous Invoice Extraction & Reconciliation Pipeline

Reviewed by Umar Abbas • Founder & Principal AI Architect

This technical reference architecture details the layout parsing, pgvector hybrid search, and post-mortem table boundary fix for autonomous invoice processing. Engineered with PostgreSQL pgvector, Reciprocal Rank Fusion (RRF), and self-hosted vLLM inference servers, the blueprint evaluates hybrid dense-sparse retrieval across complex financial billing formats.

Architecture PatternPostgreSQL pgvector Hybrid Search + RRF
Primary Constraint SolvedReliable line-item parsing without table column bleeding
StackPostgreSQL pgvector, vLLM, Python, Pydantic
Data BasisSynthetic invoice dataset and public document benchmarks
1. Project Profile & Architectural Context

Financial Document Processing Bottlenecks

Architecture Note: This reference architecture documents an engineering system designed and benchmarked internally by Esaholic engineers to validate enterprise financial document processing. Accounting workflows face severe bottlenecks: manual entry of multi-page supplier invoices takes over 14 minutes per document, introduces line-item error rates, and incurs high labor overhead.

2. System Architecture

End-to-End Extraction Pipeline

The pipeline combines layout-aware PDF tokenization, PostgreSQL pgvector hybrid search with Reciprocal Rank Fusion (RRF), and self-hosted vLLM inference servers to parse structured accounting fields.

Invoice Processing Architecture Pipeline

Interactive Flow Diagram
Invoice Processing Architecture Pipeline Data flow across PDF ingestion, layout OCR, pgvector hybrid index, vLLM JSON extraction, and ERP ledger sync. 1. Ingestion FastAPI S3 Event 2. Hybrid Search Postgres pgvector 3. LLM Extraction vLLM (Llama 3.3) 4. Rule Validation Pydantic Guard 5. ERP Sync Zero Data Retention
Stage 1: 1. Ingestion Ingestion < 12ms

Receives multi-page PDF invoice, validates checksum, and extracts layout bounding boxes.

Data flow across PDF ingestion, layout OCR, pgvector hybrid index, vLLM JSON extraction, and ERP ledger sync.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Ingestion Receives multi-page PDF invoice, validates checksum, and extracts layout bounding boxes. Ingestion < 12ms
2 2. Hybrid Search Executes parallel HNSW vector search and BM25 text match fused via RRF. Hybrid query
3 3. LLM Extraction Extracts structured accounting fields into validated Pydantic JSON. Inference serving
4 4. Rule Validation Verifies subtotal math (line_items_sum == total_amount) prior to ledger insertion. Schema check
5 5. ERP Sync Commits validated JSON transaction to ERP database and purges volatile RAM cache. Memory flush
3. Process Transformation

Operational Workflow Shift

Automating invoice parsing through deterministic schema validation and hybrid vector search replaces slow, error-prone manual keying with an automated ledger sync pipeline.

Automated Pipeline Stages
  • Automated Ingestion: Webhook events trigger instant PDF parsing and bounding-box layout segmentation.
  • Structured Extraction: vLLM inference engine extracts line items with Pydantic type safety validation.
  • Ledger Auto-Sync: Validated transactions post to general ledgers with zero human intervention.
4. Architecture Phases

Implementation Phases

The reference architecture was developed across four distinct engineering phases, moving from synthetic dataset curation to full load evaluation.

Development Phases
  • Phase 1 (Dataset & Parsing): Assembled synthetic test PDF invoices and constructed layout-aware PDF parser.
  • Phase 2 (Hybrid Indexing): Deployed PostgreSQL pgvector with HNSW vector index and BM25 tsvector keyword search.
  • Phase 3 (Inference Serving): Deployed vLLM container cluster hosting quantized models for structured JSON generation.
  • Phase 4 (Post-Mortem Audit): Conducted stress testing and resolved table boundary hallucination edge cases.
5. Mandatory Post-Mortem

What Went Wrong and How We Fixed It

Every complex AI engineering deployment encounters failure modes during stress testing. Documenting these failure points is essential for production transparency.

What Went Wrong: Table Boundary Hallucinations

During initial load testing, pure dense vector embedding models misclassified numeric table columns whenever document layout gridlines were missing. Dense vector search alone returned context from adjacent columns, causing line-item extraction mismatches on complex multi-column tables.

How We Fixed It: Layout Coordinates & RRF

We implemented two architectural fixes: (1) Added explicit 2D bounding-box coordinates to chunk metadata, and (2) Replaced pure cosine distance with Reciprocal Rank Fusion (RRF) combining HNSW dense vectors with BM25 sparse text indices in PostgreSQL. This resolved column bleeding and stabilized line-item extraction.

6. Technical Evaluation Matrix

Pure Dense Search vs. Hybrid pgvector + RRF

Engineering evaluation comparing pure dense vector search against the hybrid pgvector + Reciprocal Rank Fusion architecture.

Evaluation ParameterPure Dense Vector SearchHybrid pgvector + RRF ArchitectureArchitectural Benefit
Numerical Column MatchingSusceptible to adjacent column bleedingFused BM25 sparse lexical scoringPreserves exact numeric line items
Spatial CoordinatesUnstructured text tokens2D bounding-box spatial metadataRetains table row alignment
Output ValidationUnchecked LLM stringsStrict Pydantic arithmetic validationGuarantees line_item_sum == total
Data BoundaryThird-party external APIsSelf-hosted vLLM inside VPCEnforces zero data retention

Note: Latency and throughput figures represent internal benchmarks conducted on synthetic invoice datasets in a local evaluation environment, not client production results.

7. Engineering Takeaways

Key Architectural Lessons

1. Hybrid Search is Essential

Dense vector search alone cannot process financial documents containing numeric table columns. Combining HNSW vectors with BM25 sparse keyword indices is necessary.

2. Parent-Child Context Chunking

Indexing small child chunks guarantees precise vector matches, while passing larger parent document blocks to LLMs prevents context truncation.

3. Zero Data Retention Endpoints

Enforcing volatile memory purging immediately after token extraction guarantees compliance with enterprise financial data privacy standards.

9. Technical Blueprint FAQ

Frequently Asked Questions

Is this blueprint based on a live client or an internal engineering build?↓

This blueprint documents an internal reference engineering architecture built and benchmarked by Esaholic engineers to validate production document extraction performance without exposing proprietary client data.

How was the invoice extraction precision evaluated?↓

Precision was evaluated across a benchmark dataset of multi-page unstructured PDF invoices compared against ground-truth manual accountant entries.

What security measures protect sensitive accounting data during processing?↓

The pipeline deploys Zero Data Retention cloud endpoints wrapped in customer-managed KMS keys, purging volatile RAM memory immediately after token extraction.

How did the post-mortem fix resolve vector search accuracy degradation?↓

Switching from pure cosine dense vector search to a Reciprocal Rank Fusion (RRF) hybrid model combining HNSW dense vectors with BM25 sparse keyword indices restored accuracy on numerical tables.

Ready to Benchmark Your Document Extraction Pipeline?

Schedule a technical architecture review with Founder & Principal AI Architect Umar Abbas to evaluate custom pgvector RAG pipelines and vLLM inference deployments under NDA.

Explore Generative AI Services →