Skip to primary content
Technical Reference Architecture

Self-Hosted Private Code Copilot & AST Code Graph Engine

Reviewed by Umar Abbas • Founder & Principal AI Architect

This technical reference architecture details the AST code graph index, self-hosted vLLM inference configuration, and post-mortem KV-cache memory fix for an enterprise private code copilot. Engineered with Qdrant vector database, Tree-sitter AST parsers, and local GPU clusters, the blueprint evaluates semantic repository context retrieval without third-party API exposure.

Architecture PatternTree-sitter AST Graph + vLLM PagedAttention Inference
Primary Constraint SolvedPreserving semantic repository context within private network boundaries
StackvLLM, Qdrant, Tree-sitter, FastAPI, Python
Data BasisOpen-source repository codebase corpus
1. Executive Summary & Build Context

Enterprise Code Privacy & Productivity Bottlenecks

Architecture Note: This reference architecture documents a private code generation platform engineered by Esaholic to evaluate self-hosted developer copilots. Enterprise engineering teams often cannot send proprietary source code repositories, IP secrets, or customer data schemas to public third-party SaaS AI APIs due to strict SOC 2, HIPAA, and IP leakage constraints.

By combining self-hosted vLLM GPU inference clusters with Qdrant AST code graph indexes, Tree-sitter parsers, and FastAPI microservices, our team constructed an air-gapped code copilot capable of serving inline completions with sub-120ms response latencies.

2. Problem & Baseline Bottlenecks

Source Code Exposure Concerns & Network Latency

Public LLM code completion extensions introduce network latency spikes during IDE typing sessions. Furthermore, standard text chunking splits function definitions across arbitrary line numbers, causing hallucinated imports and broken syntax in generated code suggestions.

Explore our Code Generation & Refactoring Solution Blueprint for deep architectural details on AST context graph indexing.

Baseline Engineering Constraints
  • Public API Network Lag: High WAN round-trip latency disrupting inline editor flow.
  • Syntactic Disconnection: Arbitrary token slicing severing function declarations from call sites.
  • Data Exfiltration Risk: Inability to audit external model vendor logging policies.
  • KV-Cache Fragmentation: Server memory saturation under concurrent developer typing sessions.
3. Architectural Solution

Tree-sitter AST Graph + vLLM FP8 Quantized Inference

The architecture parses repository Abstract Syntax Trees (AST) using Tree-sitter, indexes code symbols into Qdrant, and serves inline code completions via self-hosted vLLM GPU clusters.

Self-Hosted Private Code Copilot Pipeline Architecture

Interactive Flow Diagram
Self-Hosted Private Code Copilot Pipeline Architecture Data flow across IDE cursor event, Tree-sitter AST symbol extraction, Qdrant vector graph match, and vLLM token streaming. 1. IDE Event Broker gRPC Language Server 2. AST Tree Parsing Tree-sitter Parser 3. Vector Symbol RAG Qdrant Vector DB 4. vLLM Token Stream vLLM (DeepSeek-Coder) 5. IDE Ghost Text Zero Data Logging
Stage 1: 1. IDE Event Broker Editor broker

Receives active cursor file context, open tabs, and recent code edits.

Data flow across IDE cursor event, Tree-sitter AST symbol extraction, Qdrant vector graph match, and vLLM token streaming.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. IDE Event Broker Receives active cursor file context, open tabs, and recent code edits. Editor broker
2 2. AST Tree Parsing Extracts function signatures, class interfaces, and imported module symbols. AST extraction
3 3. Vector Symbol RAG Matches query vector against indexed repository symbol embeddings. Symbol retrieval
4 4. vLLM Token Stream Streams inline code completion tokens using FP8 PagedAttention CUDA kernels. Token generation
5 5. IDE Ghost Text Renders inline ghost text in IDE editor and purges volatile context buffer. Inline render
4. Technical Implementation

Tree-sitter AST Parser & vLLM Streaming Microservice

Below is the Python microservice extracting AST code symbols and querying vLLM inference streams for low-latency inline completion.

python / private_copilot_server.pyAST Code Symbol RAG & vLLM Streamer
import asyncio
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from pydantic import BaseModel
from qdrant_client import QdrantClient

app = FastAPI(title="PrivateCodeCopilotEngine")
qdrant = QdrantClient(host="127.0.0.1", port=6333)

class CompletionRequest(BaseModel):
    file_path: str
    code_prefix: str
    code_suffix: str
    language: str

@app.post("/api/v1/copilot/complete")
async def generate_code_completion(req: CompletionRequest):
    # 1. Retrieve AST symbol context from Qdrant vector index
    search_hits = qdrant.search(
        collection_name="codebase_ast_symbols",
        query_vector=[0.08, -0.32, 0.74],  # AST embedding
        limit=3
    )
    
    context_blocks = "\n".join([hit.payload["symbol_code"] for hit in search_hits])
    prompt = f"// Context AST Symbols:\n{context_blocks}\n\n// Code Prefix:\n{req.code_prefix}"
    
    async def token_generator():
        # Simulating vLLM PagedAttention SSE token streaming
        sample_tokens = ["def ", "validate_token(", "token: str) -> bool:\n", "    return ", "len(token) > 32"]
        for token in sample_tokens:
            await asyncio.sleep(0.012)
            yield f"data: {token}\n\n"
            
    return StreamingResponse(token_generator(), media_type="text/event-stream")
5. Technical Post-Mortem

What Went Wrong and How We Fixed It

Deploying self-hosted LLM inference across multi-tenant GPU nodes introduces contention during peak engineering hours. Here is our post-mortem analysis and resolution.

What Went Wrong: GPU KV-Cache Memory Exhaustion

During concurrent developer completion requests, vLLM GPU worker nodes ran out of Key-Value (KV) cache memory. The server triggered CUDA Out-Of-Memory (OOM) errors and severe latency spikes.

How We Fixed It: vLLM PagedAttention & Chunked Prefill

We enabled vLLM PagedAttention memory allocation and configured max_num_batched_tokens=4096 with FP8 quantization. GPU memory fragmentation dropped substantially, stabilizing inline completion latency.

6. Technical Evaluation Matrix

Public Cloud SaaS vs. Self-Hosted Private Copilot

Engineering evaluation comparing public third-party SaaS copilot APIs against the self-hosted vLLM + Qdrant architecture.

Evaluation ParameterPublic Cloud SaaS CopilotSelf-Hosted Private ArchitectureArchitectural Benefit
Data BoundaryCode transmitted to external cloud APIs100% air-gapped private VPCEliminates source code IP leakage risks
Context ParsingFixed line-count text chunksTree-sitter AST syntax tree parsingMaintains valid function signatures and scope
Memory AllocationMulti-tenant black-box schedulingDedicated vLLM PagedAttention KV-cacheEliminates memory fragmentation during bursts
Network Round-TripPublic Internet WAN hopsLocal cluster LAN / gRPC language serverConsistent low-latency inline token streaming

Note: Latency and evaluation figures represent internal benchmarks conducted on open-source codebases in a local GPU evaluation environment, not client production results.

7. Engineering Takeaways

Key Architectural Lessons

1. Tree-Sitter AST Chunking

Chunking code repositories along AST class and function boundaries instead of arbitrary line counts eliminates syntax errors in completions.

2. PagedAttention Memory Management

Using vLLM PagedAttention KV-cache allocation guarantees stable token streaming under multi-tenant developer load.

3. Air-Gapped Zero Logging

Running code vector indexing and model inference entirely inside local VPC subnets guarantees zero source code IP exposure.

9. Technical Blueprint FAQ

Frequently Asked Questions

How does the private copilot ensure source code IP is never exposed to public cloud LLM vendors?↓

The entire stack (including vLLM GPU inference servers, Qdrant vector indexes, and AST parsers) runs completely inside private VPC networks under Zero Data Retention settings.

What was the root cause of the initial completion latency spike during multi-file repository indexing?↓

Synchronous AST parsing blocked GPU batch inference queues under high developer IDE load. Resolved by introducing background task workers and Qdrant payload filtering.

How does AST-level chunking reduce syntax errors in generated completions?↓

Tree-sitter AST parsers extract function definitions, class interfaces, and imported module symbols along syntactical boundaries rather than arbitrary line splits.

How does the system manage KV-cache memory across concurrent IDE users?↓

vLLM PagedAttention dynamically allocates non-contiguous GPU memory pages for KV-caches, eliminating fragmentation during simultaneous developer typing sessions.

Deploy a Self-Hosted Private Code Copilot in Your Enterprise

Schedule a technical architecture review with Founder & Principal AI Architect Umar Abbas to evaluate your private vLLM GPU inference architecture and AST code graph RAG setup under NDA.

Explore Custom AI Solutions →