Skip to primary content
Engineering Leadership • Full-Time • Remote (UK / Europe ±3h) • Posted: August 14, 2026

Principal RAG Systems Architect (pgvector / vLLM)

As Principal RAG Systems Architect at Esaholic, you will lead the technical design and benchmarking of enterprise Retrieval-Augmented Generation (RAG) pipelines. You will own PostgreSQL pgvector HNSW indexing, Reciprocal Rank Fusion (RRF) search algorithms, Cohere cross-encoder rerankers, and high-throughput vLLM inference clusters.

Base Salary Band £110k - £145k GBP
Vector Engine pgvector / Qdrant
Serving Runtime vLLM (FP8 / AWQ)
Location Policy Remote (UK/EU)
Compensation Note: Target base range £110,000 to £145,000 + equity options.
Leadership Execution Plan

What You Will Accomplish in Your First 90 Days

As our Principal RAG Architect, you will establish the technical standards for high-precision retrieval across all enterprise client projects.

Days 1 – 30: Retrieval Audit

Benchmark Production Hybrid Search

  • Audit our existing `pgvector` HNSW index configurations and tuning parameters (`m`, `ef_construction`).
  • Benchmark top-3 recall accuracy comparing bi-encoder cosine search against Reciprocal Rank Fusion (RRF).
  • Optimize sentence transformer embedding batching queues for sub-30ms latency.
Days 31 – 60: Architecture Ownership

vLLM Inference & Reranking Subsystem

  • Architect high-concurrency vLLM serving clusters across NVIDIA A100/H100 GPU nodes with FP8 quantization.
  • Integrate Cohere Rerank v3 cross-encoder models to filter top-30 retrieval candidate lists.
  • Design parent-child document chunking pipelines preserving 2D table bounding-box metadata.
Days 61 – 90: Enterprise Scaling

Multi-Tenant Vector Scaling

  • Deploy multi-tenant vector isolation strategies for 100M+ document enterprise search indices.
  • Establish automated evaluation frameworks (Ragas / TruLens) tracking faithfulness and context recall.
  • Mentor senior AI engineers on vector index memory sizing and SIMD hardware acceleration.
Stack & Leadership Domain

Systems & Technologies You Will Own

Primary Vector & Search Systems
  • PostgreSQL pgvector Core: HNSW index tuning, `tsvector` BM25 full-text integration, RRF SQL query fusion.
  • Cross-Encoder Reranking Engine: Cohere Rerank v3 and `bge-reranker-large` joint-attention pipelines.
  • vLLM GPU Inference Cluster: Containerized Open-Weights LLM serving (Llama 3.3 70B FP8, Continuous Batching, PagedAttention).
  • Document Processing Engine: Parent-child chunkers, layout-aware OCR, 2D bounding-box parsers.
Named Tooling & Infrastructure Stack
  • Languages & Runtimes: Python 3.12+, PyTorch 2.4+, CUDA 12, FastAPI, AsyncIO.
  • Vector Storage Engines: PostgreSQL 16 (`pgvector`), Qdrant, Pinecone.
  • Evaluation & Observability: Ragas, TruLens, LangSmith, Prometheus, Grafana.
  • Cloud Infrastructure: AWS (EC2 GPU instances), Kubernetes, Docker, Terraform.
Realistic Boundaries

What We Do NOT Expect from You

1. Theoretical AI Math Isolation

We do not expect pure academic research papers. We build production systems for enterprise clients - we care about sub-50ms p95 latencies, RAM sizing, and index build SLAs.

2. Synthetic LeetCode Puzzle Tests

We will never ask you to solve dynamic programming puzzles under a 45-minute countdown clock. We evaluate real vector architecture design and SQL performance tuning.

3. Uncontrolled Overtime Culture

We respect your life outside of work. Engineering sprints are planned realistically, with automated staging tests preventing late-night deployment panics.

4. Pure Management Overhead

This is a hands-on technical leadership role (approx 70% technical design & code, 30% mentorship). You will not be bogged down in administrative bureaucracy.

Transparent Interview Loop

How We Interview Candidates

Our hiring process is transparent, thorough, and completed within 14 calendar days.

Principal RAG Architect Hiring Process Roadmap

Phase Delivery Roadmap
Phase 1 30 Mins (Call)

Architect Intro

Conversation with technical leadership discussing your retrieval engineering background and system philosophy.

Deliverables:
  • ✓ Mutual Alignment Check
  • ✓ Role & Tech Review
Phase 2 60 Mins (Video)

RAG System Design

Practical vector architecture session. We present a 10M-document hybrid search scenario and design the SQL schema & reranking pipeline together.

Deliverables:
  • ✓ Hybrid Search Architecture
  • ✓ vLLM Benchmark Strategy
Phase 3 45 Mins (Video)

Architecture Leadership Sync

Technical conversation with Umar Abbas (Founder & Principal AI Architect) covering GPU memory allocation, HNSW tuning, and team mentorship.

Deliverables:
  • ✓ Deep Systems Sync
  • ✓ Team Culture Check
Phase 4 48 Hours

Offer & Decision

Formal offer letter detailing base salary band, equity option package, equipment budget, and start date options.

Deliverables:
  • ✓ Signed Offer Letter
  • ✓ Onboarding Roadmap
Four-step hiring pipeline designed to evaluate practical retrieval engineering leadership within 14 days.
Text alternative for screen readers & search engines
  1. Phase 1: Architect Intro (30 Mins (Call)) - Conversation with technical leadership discussing your retrieval engineering background and system philosophy. Key deliverables: Mutual Alignment Check, Role & Tech Review.
  2. Phase 2: RAG System Design (60 Mins (Video)) - Practical vector architecture session. We present a 10M-document hybrid search scenario and design the SQL schema & reranking pipeline together. Key deliverables: Hybrid Search Architecture, vLLM Benchmark Strategy.
  3. Phase 3: Architecture Leadership Sync (45 Mins (Video)) - Technical conversation with Umar Abbas (Founder & Principal AI Architect) covering GPU memory allocation, HNSW tuning, and team mentorship. Key deliverables: Deep Systems Sync, Team Culture Check.
  4. Phase 4: Offer & Decision (48 Hours) - Formal offer letter detailing base salary band, equity option package, equipment budget, and start date options. Key deliverables: Signed Offer Letter, Onboarding Roadmap.

Compensation & Benefits Package

  • Base Salary: £110,000 to £145,000 GBP per annum (based on technical experience).
  • Equity Options: Senior equity option package in Esaholic.
  • Remote Work Setup: £2,500 workstation budget (MacBook Pro M3 Max / dual 4K displays).
  • Learning & Research Budget: £2,000 annual budget for AI conferences and technical research papers.
  • Time Off: 30 business days paid annual leave + UK bank holidays.

Location & Work Policy

This role is Remote-First for candidates residing in the United Kingdom or European time zones (GMT ± 3 hours). Candidates in London have optional access to our hybrid office space.

How to Apply

Send an email to our engineering team with your CV and links to your GitHub profile or published technical system architecture notes:

careers@esaholic.com

Subject Line: Application: Principal RAG Architect - [Your Name]

Include: CV PDF, GitHub / Architecture link, and a 2-sentence note on your most complex vector or database optimization project.