Principal RAG Systems Architect (pgvector / vLLM)
As Principal RAG Systems Architect at Esaholic, you will lead the technical design and benchmarking of enterprise Retrieval-Augmented Generation (RAG) pipelines. You will own PostgreSQL pgvector HNSW indexing, Reciprocal Rank Fusion (RRF) search algorithms, Cohere cross-encoder rerankers, and high-throughput vLLM inference clusters.
What You Will Accomplish in Your First 90 Days
As our Principal RAG Architect, you will establish the technical standards for high-precision retrieval across all enterprise client projects.
Benchmark Production Hybrid Search
- Audit our existing `pgvector` HNSW index configurations and tuning parameters (`m`, `ef_construction`).
- Benchmark top-3 recall accuracy comparing bi-encoder cosine search against Reciprocal Rank Fusion (RRF).
- Optimize sentence transformer embedding batching queues for sub-30ms latency.
vLLM Inference & Reranking Subsystem
- Architect high-concurrency vLLM serving clusters across NVIDIA A100/H100 GPU nodes with FP8 quantization.
- Integrate Cohere Rerank v3 cross-encoder models to filter top-30 retrieval candidate lists.
- Design parent-child document chunking pipelines preserving 2D table bounding-box metadata.
Multi-Tenant Vector Scaling
- Deploy multi-tenant vector isolation strategies for 100M+ document enterprise search indices.
- Establish automated evaluation frameworks (Ragas / TruLens) tracking faithfulness and context recall.
- Mentor senior AI engineers on vector index memory sizing and SIMD hardware acceleration.
Systems & Technologies You Will Own
- PostgreSQL pgvector Core: HNSW index tuning, `tsvector` BM25 full-text integration, RRF SQL query fusion.
- Cross-Encoder Reranking Engine: Cohere Rerank v3 and `bge-reranker-large` joint-attention pipelines.
- vLLM GPU Inference Cluster: Containerized Open-Weights LLM serving (Llama 3.3 70B FP8, Continuous Batching, PagedAttention).
- Document Processing Engine: Parent-child chunkers, layout-aware OCR, 2D bounding-box parsers.
- Languages & Runtimes: Python 3.12+, PyTorch 2.4+, CUDA 12, FastAPI, AsyncIO.
- Vector Storage Engines: PostgreSQL 16 (`pgvector`), Qdrant, Pinecone.
- Evaluation & Observability: Ragas, TruLens, LangSmith, Prometheus, Grafana.
- Cloud Infrastructure: AWS (EC2 GPU instances), Kubernetes, Docker, Terraform.
What We Do NOT Expect from You
We do not expect pure academic research papers. We build production systems for enterprise clients - we care about sub-50ms p95 latencies, RAM sizing, and index build SLAs.
We will never ask you to solve dynamic programming puzzles under a 45-minute countdown clock. We evaluate real vector architecture design and SQL performance tuning.
We respect your life outside of work. Engineering sprints are planned realistically, with automated staging tests preventing late-night deployment panics.
This is a hands-on technical leadership role (approx 70% technical design & code, 30% mentorship). You will not be bogged down in administrative bureaucracy.
How We Interview Candidates
Our hiring process is transparent, thorough, and completed within 14 calendar days.
Principal RAG Architect Hiring Process Roadmap
Phase Delivery RoadmapArchitect Intro
Conversation with technical leadership discussing your retrieval engineering background and system philosophy.
- ✓ Mutual Alignment Check
- ✓ Role & Tech Review
RAG System Design
Practical vector architecture session. We present a 10M-document hybrid search scenario and design the SQL schema & reranking pipeline together.
- ✓ Hybrid Search Architecture
- ✓ vLLM Benchmark Strategy
Architecture Leadership Sync
Technical conversation with Umar Abbas (Founder & Principal AI Architect) covering GPU memory allocation, HNSW tuning, and team mentorship.
- ✓ Deep Systems Sync
- ✓ Team Culture Check
Offer & Decision
Formal offer letter detailing base salary band, equity option package, equipment budget, and start date options.
- ✓ Signed Offer Letter
- ✓ Onboarding Roadmap
Text alternative for screen readers & search engines
- Phase 1: Architect Intro (30 Mins (Call)) - Conversation with technical leadership discussing your retrieval engineering background and system philosophy. Key deliverables: Mutual Alignment Check, Role & Tech Review.
- Phase 2: RAG System Design (60 Mins (Video)) - Practical vector architecture session. We present a 10M-document hybrid search scenario and design the SQL schema & reranking pipeline together. Key deliverables: Hybrid Search Architecture, vLLM Benchmark Strategy.
- Phase 3: Architecture Leadership Sync (45 Mins (Video)) - Technical conversation with Umar Abbas (Founder & Principal AI Architect) covering GPU memory allocation, HNSW tuning, and team mentorship. Key deliverables: Deep Systems Sync, Team Culture Check.
- Phase 4: Offer & Decision (48 Hours) - Formal offer letter detailing base salary band, equity option package, equipment budget, and start date options. Key deliverables: Signed Offer Letter, Onboarding Roadmap.
Compensation & Benefits Package
- Base Salary: £110,000 to £145,000 GBP per annum (based on technical experience).
- Equity Options: Senior equity option package in Esaholic.
- Remote Work Setup: £2,500 workstation budget (MacBook Pro M3 Max / dual 4K displays).
- Learning & Research Budget: £2,000 annual budget for AI conferences and technical research papers.
- Time Off: 30 business days paid annual leave + UK bank holidays.
Location & Work Policy
This role is Remote-First for candidates residing in the United Kingdom or European time zones (GMT ± 3 hours). Candidates in London have optional access to our hybrid office space.
How to Apply
Send an email to our engineering team with your CV and links to your GitHub profile or published technical system architecture notes:
Subject Line: Application: Principal RAG Architect - [Your Name]
Include: CV PDF, GitHub / Architecture link, and a 2-sentence note on your most complex vector or database optimization project.