Project Overview
| Client Industry | Finance / Insurance |
| Business Type | Regional commercial insurance provider with ~150 underwriters, claims adjusters, and compliance staff |
| Project Duration | 12 weeks |
| AI Service Provided | LLM and RAG Development |
| Technologies Used | OpenAI GPT-4 class models, LlamaIndex, Pinecone (vector database), Python, FastAPI, PostgreSQL, Microsoft Azure, Okta SSO |
The Client Challenge
The client’s underwriters and claims staff needed to reference internal policy wording, state-specific regulatory guidelines, and internal underwriting manuals constantly but that knowledge was scattered across thousands of PDFs, SharePoint folders, and legacy document management systems.
The operational impact:
- Underwriters spent an estimated 6–8 hours a week searching for specific policy language or coverage exceptions, often resorting to asking senior colleagues instead of searching documents directly.
- Answers varied depending on who you asked inconsistent interpretation of policy wording created compliance risk and rework.
- New hires took months to become productive because so much knowledge lived in people’s heads rather than searchable systems.
- Compliance staff needed to verify regulatory requirements across 14 states, each with different filing and coverage rules, and cross-referencing them manually was slow and error-prone.
- Existing keyword search tools returned document lists, not answers staff still had to read through long PDFs to find the specific clause they needed.
The client had experimented with a general-purpose AI chatbot, but it either refused to answer without document access or, worse, generated plausible-sounding but incorrect policy interpretations an unacceptable risk in a regulated industry.
Our Solution
Air Brite Labs built an internal knowledge assistant grounded entirely in the client’s own policy documents, underwriting manuals, and state regulatory filings using retrieval-augmented generation (RAG) ensuring every answer is traceable back to a real source document.
Staff ask questions in plain language (“What’s the wind/hail deductible exception for commercial property in Texas?”) and receive a direct answer with the exact source document and page cited, not just a document list.
Key capabilities:
- Grounded, cited answers — every response includes a direct citation to the source clause, so staff can verify accuracy instantly rather than trusting the AI blindly.
- State-aware retrieval — the system understands regulatory jurisdiction context and retrieves the correct state-specific documents rather than mixing rules across states.
- Role-based access — underwriters, claims, and compliance see the document sets relevant to their role, respecting existing internal access permissions.
- Refusal on uncertainty — when the system can’t find a confident match in the source documents, it says so explicitly rather than guessing.
- Query history and feedback loop — staff can flag inaccurate or unhelpful answers, feeding directly into ongoing tuning.
Technical Approach
LLM Layer: OpenAI GPT-4 class models for answer generation, constrained to only synthesize from retrieved document chunks the model is explicitly instructed not to draw on general knowledge for policy-specific questions.
RAG Architecture: Documents are chunked with LlamaIndex using a hierarchy-aware splitting strategy that preserves clause structure (so a coverage exception isn’t split mid-sentence from its qualifying condition). Chunks are embedded and stored in Pinecone, with metadata tagging for document type, state jurisdiction, and effective date.
Retrieval Logic: Query-time retrieval combines semantic search with metadata filtering — a query about Texas commercial property first filters to Texas + commercial property documents, then ranks by semantic relevance within that filtered set, significantly reducing cross-jurisdiction errors.
Data Processing: An ingestion pipeline handles the client’s existing document formats (PDF, Word, scanned documents requiring OCR), with a versioning system that flags outdated documents when policy manuals are updated, preventing the system from citing superseded rules.
Integrations: Okta SSO for authentication and role-based access control; internal SharePoint connector for automatic ingestion of newly published documents.
Security & Scalability: Deployed on Microsoft Azure to align with the client’s existing compliance and data residency requirements. All document access respects existing role permissions the RAG system never surfaces content a user wouldn’t otherwise have access to. The vector index scales independently of the application layer, so adding new document sets doesn’t require re-architecting retrieval.
Implementation Process
- Discovery & Requirement Analysis (Weeks 1–2): Audited document sources, mapped access permission requirements, and identified the highest-friction query types from underwriters and compliance staff.
- Prototype / PoC (Weeks 3–4): Built a working RAG pipeline against a subset of Texas and California policy documents, validated citation accuracy with senior underwriters.
- Development (Weeks 5–8): Scaled ingestion to the full 12,000-page document set across all 14 states, built the state-aware retrieval filtering and role-based access layer.
- Integration (Week 9): Connected Okta SSO and the SharePoint ingestion pipeline for ongoing document updates.
- Testing (Week 10): Ran a structured accuracy evaluation with 200 real underwriting questions, scored by senior staff against known-correct answers.
- Deployment (Week 11): Rolled out to a pilot group of 20 underwriters before company-wide release.
- Optimization (Week 12 and ongoing): Refined chunking strategy based on retrieval misses identified during pilot, added the refusal-on-uncertainty behavior after observing edge cases.
Key Features Delivered
- RAG-powered knowledge assistant grounded in 12,000+ pages of internal policy and regulatory documents
- Source-cited answers with document and page-level traceability
- State-aware retrieval filtering across 14 jurisdictions
- Role-based access control aligned to existing internal permissions
- Explicit refusal behavior when confident answers aren’t available
- Automated document ingestion pipeline with version tracking
- Feedback and query history logging for continuous improvement
Business Results
- Average time to find a specific policy answer dropped from 15–20 minutes to under 90 seconds
- Underwriter time spent searching for information reduced by an estimated 5+ hours per person per week
- New hire ramp-up time for policy familiarity shortened by approximately 4–6 weeks, based on manager feedback
- Citation accuracy verified at 94% in structured evaluation against known-correct answers, with all remaining cases correctly flagged as low-confidence rather than answered incorrectly
- Compliance team reported faster, more consistent cross-state rule verification, reducing manual cross-referencing work
- Query volume grew organically post-launch as staff shifted from asking colleagues to querying the system directly
Technology Stack
| Layer | Technology |
|---|---|
| LLM | OpenAI GPT-4 class models |
| RAG Framework | LlamaIndex |
| Vector Database | Pinecone |
| Backend | Python, FastAPI |
| Database | PostgreSQL |
| Cloud Infrastructure | Microsoft Azure |
| Authentication | Okta SSO |
| Document Source | SharePoint connector |
Why the Solution Worked
Grounding every answer in cited source documents, and explicitly refusing to answer when retrieval confidence was low, was the difference between a useful tool and a compliance liability. In a regulated industry, trust depends on verifiability not just accuracy.
The state-aware retrieval filtering was equally critical. Generic semantic search alone would have occasionally surfaced the right clause from the wrong state, which is a subtle but serious error in insurance. Building jurisdiction awareness directly into the retrieval logic, rather than relying on the LLM to sort it out, closed that gap.
Future Scalability
The RAG architecture is designed to extend well beyond its initial scope. The client is currently planning:
- Expanding document coverage to include historical claims files for claims adjuster support
- Adding a drafting-assist mode that helps underwriters write policy endorsements using verified language
- Extending role-based access to external broker partners for self-service policy lookups
- Integrating with the client’s policy administration system for real-time document updates rather than batch ingestion
Because ingestion, retrieval, and access control are built as separate, modular layers, each of these expansions can be added incrementally without disrupting the live system.
Final Outcome
What was once a scattered, tribal-knowledge-dependent process is now a fast, verifiable, and auditable knowledge system. Underwriters and compliance staff get accurate answers in seconds instead of minutes, backed by real source citations reducing both operational friction and regulatory risk.



