> ## Documentation Index
> Fetch the complete documentation index at: https://docs.uselayerup.com/llms.txt
> Use this file to discover all available pages before exploring further.

# IDI Underwriting Agent

> End-to-end architecture, document intelligence pipeline, confidence calibration, and performance benchmarks for the Individual Disability Income underwriting workflow.

# IDI Underwriting Agent — architecture, benchmarks & workflow deep dive.

The Individual Disability Income (IDI) Underwriting Agent is a fully sovereign, AOP-governed reasoning workload purpose-built for new business underwriting across individual disability income product lines. It ingests the full application packet — application form, Attending Physician Statements, occupation duty questionnaires, financial documents, and third-party data feeds — processes every document simultaneously in a unified reasoning context, and produces a structured, evidence-cited, confidence-scored recommendation within a single automated pass.

This page covers the IDI-specific architecture, the insurance ontology embedded in the Agent Operating Procedure, how the confidence engine is calibrated and validated for this workflow, independently measured benchmarks against the IDI-specific synthetic validation set, and the KPI impact model for underwriting operations.

| attribute          | value                                                                                                                                |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------ |
| workflow           | Individual Disability Income — New Business Underwriting                                                                             |
| case types         | Non-cancelable · Guaranteed renewable · Business overhead expense                                                                    |
| document intake    | Application form · APS · Attending Physician Statements · Rx history · Financials · Occupation duty questionnaire · IREX · MIB · MVR |
| decision outputs   | Approve as Applied · Approve with Exclusion · Approve Modified · Defer Pending Requirements · Escalate · Decline                     |
| confidence scoring | Composite 0–100 · Five independent signal domains · IDI-calibrated weight set                                                        |
| deployment model   | Private VPC / VNet — your cloud account, your network                                                                                |

***

## 1 — IDI-specific Agent Operating Procedure

The AOP is the agent's underwriting contract. For the IDI workflow, it is not a generic document-processing instruction set — it is built on an **insurance-native ontology** that Layerup has assembled from first-principles IDI underwriting practice. Your enterprise SOP is layered on top of this baseline, not substituted for it.

The IDI AOP schema covers three primary domains that are specific to disability income underwriting and do not exist in Layerup's other workflow agents.

***

### 1.1 Occupation classification schema

Occupation analysis is the primary risk driver in IDI underwriting. The AOP encodes a structured occupation intelligence schema that governs how the agent evaluates every occupational dimension of the application.

<CardGroup cols={2}>
  <Card title="Occupation Class Definitions" icon="briefcase">
    The AOP contains the full occupation class definition table for your product line — including class boundaries, duty mix requirements by class, supervisory vs. manual split thresholds, and any product-specific class restrictions. The agent evaluates the applicant's submitted job title and duty questionnaire against this table in every case.
  </Card>

  <Card title="Title vs. Duties Mismatch Detection" icon="arrows-left-right">
    The agent independently evaluates the submitted occupation title and the
    described duties. Where the duty profile described in the questionnaire or APS
    is inconsistent with the submitted occupation class, the mismatch is flagged
    as a first-class underwriting finding — not buried in a generic inconsistency
    flag.
  </Card>

  <Card title="Duty Mix Quantification" icon="chart-pie">
    The AOP defines the duty mix breakpoints — the percentage split between
    manual, supervisory, and clerical functions — that determine class eligibility
    for your product. The agent attempts to resolve the duty mix from all
    available sources (questionnaire, APS narrative, employer letter) and surfaces
    it as a structured output field alongside a confidence sub-score for the
    resolution.
  </Card>

  <Card title="Hazardous Activity & Avocation Flags" icon="triangle-exclamation">
    The AOP encodes a curated list of avocations and hazardous activities that
    affect IDI risk classification or eligibility. The agent cross-references the
    application form, APS narrative, and MIB record against this list and surfaces
    any matches as occupation-domain findings with the specific source document
    cited.
  </Card>

  <Card title="Specialty & Licensure Verification" icon="graduation-cap">
    For professional applicant classes (physicians, attorneys, CPAs, and
    equivalent), the AOP specifies the specialty risk tables and any
    licensure-based eligibility criteria relevant to your product. The agent
    checks the stated specialty against the APS authorship credentials and any
    available third-party professional database integration.
  </Card>

  <Card title="Multi-Employer & Contingent Income" icon="building">
    The AOP includes structured handling for applicants with non-standard employment arrangements — multiple employers, partnership income, self-employment, contract work, and deferred compensation. Each arrangement type has its own documentation sufficiency requirements and income verification logic encoded in the schema.
  </Card>
</CardGroup>

***

### 1.2 Financial documentation requirements

IDI benefit amounts are directly tied to earned income, making financial verification a core underwriting dimension. The AOP encodes income verification requirements by income type, documentation hierarchy, and cross-reference logic.

| income type                | primary document                              | secondary document                      | cross-reference logic                                                                                  |
| -------------------------- | --------------------------------------------- | --------------------------------------- | ------------------------------------------------------------------------------------------------------ |
| W-2 employment             | Most recent two years W-2                     | Employer letter (if income > threshold) | YTD pay stub reconciliation against annualised W-2                                                     |
| Self-employed / Schedule C | Two years federal tax returns (all schedules) | CPA letter or P\&L statement            | Trend analysis: year-over-year variance flagged if > AOP-defined tolerance                             |
| Partnership / S-Corp       | K-1 for two years + corporate returns         | CPA letter                              | Active participation verification; passive income exclusion applied                                    |
| Physician / Professional   | Two years personal returns + business returns | Specialty income benchmarks             | Gross vs. net income reconciliation; overhead exclusion for BOE cases                                  |
| Deferred / Bonus income    | Base + prior two years bonus schedule         | Board resolution or offer letter        | Bonus included at AOP-defined inclusion percentage; deferred excluded from monthly benefit calculation |
| Fluctuating / Commission   | Three years average, personal returns         | YTD commission statement                | Trend direction applied; declining trend triggers additional financial requirement                     |

The agent does not compute a binary pass/fail against these requirements. It produces an **income sufficiency score** for each applicable income type, identifies which documents are present and which are absent, and generates a prioritized requirements list ranked by the evidentiary gap's impact on the confidence composite.

***

### 1.3 Medical flagging criteria

The IDI AOP contains a structured medical flagging schema that is separate from the generic document extraction layer. Medical evidence is evaluated in two passes: first for extraction quality and completeness, then for underwriting significance under the encoded IDI medical guidelines.

<CardGroup cols={2}>
  <Card title="Condition Category Classification" icon="heart-pulse">
    Medical history extracted from the APS, Rx history, and IREX record is classified against the AOP's condition category table. Categories include musculoskeletal, mental health and nervous system, cardiovascular, oncology, autoimmune, and metabolic — each with distinct IDI underwriting significance weights and AOP-encoded flagging thresholds.
  </Card>

  <Card title="Prescription History Analysis" icon="capsules">
    The Rx history feed is normalized and cross-referenced against the AOP's
    medication significance table. Medications not disclosed on the application
    that appear in the Rx record are flagged as cross-evidence inconsistencies.
    Drug class patterns indicating undisclosed conditions are surfaced as APS
    requirement triggers.
  </Card>

  <Card title="Recency & Chronicity Scoring" icon="calendar">
    Conditions are evaluated not only for type but for recency and treatment
    pattern. The AOP encodes condition recency windows (e.g., conditions with last
    treatment within 24 months receive different treatment than resolved
    conditions outside the window) and chronic condition criteria that affect
    exclusion rider eligibility.
  </Card>

  <Card title="Mental Health & Substance Flags" icon="brain">
    The AOP includes specific flagging logic for mental health history and
    substance use, consistent with your product's underwriting guidelines and
    applicable state regulatory requirements. These findings are surfaced with
    their APS source citations and never inferred from Rx data alone without
    corroborating documentation.
  </Card>

  <Card title="APS Adequacy Assessment" icon="file-medical">
    The agent evaluates whether the APS(es) received are adequate to underwrite
    the disclosed conditions — checking physician specialty alignment, treatment
    recency coverage, and completeness of history. Where the APS is inadequate,
    the agent generates a targeted APS requirement specifying exactly what the
    follow-up must address.
  </Card>

  <Card title="Intra-Medical Timeline Validation" icon="timeline">
    Treatment dates, prescription fill dates, diagnostic dates, and the narrative timeline across all medical documents are cross-referenced for coherence. Timeline inconsistencies — e.g., a prescription filled for a condition the APS shows was not diagnosed until a later date — are surfaced as first-class medical evidence flags.
  </Card>
</CardGroup>

***

## 2 — IDI workflow architecture

The following diagram describes the complete data and reasoning flow for a single IDI underwriting case. Every component executes within your cloud account boundary.

```mermaid theme={null}
flowchart TD
  subgraph intake ["Intake Layer — Your Systems"]
    APP[Application Form<br/>+ Duty Questionnaire]
    MED[Medical Records<br/>APS · IREX]
    RX[Rx History<br/>Pharmacy Claims DB]
    FIN[Financial Documents<br/>Tax Returns · W-2 · Pay Stubs]
    MIB[MIB Record]
    MVR[Motor Vehicle Report]
    S3IN[S3 Input Bucket<br/>CMK-encrypted]
  end

  subgraph ingestion ["Document Intelligence Layer — Private Subnet"]
    OCR[Multi-pass OCR Pipeline<br/>Scanned PDFs · Handwritten Forms · Degraded Docs]
    PARSE[Document Parser<br/>Layout · Tables · Handwriting · Signatures]
    NORM[Schema Normalisation<br/>IDI Case Schema]
    FLAG[Ambiguity Flagger<br/>Low-confidence fields marked]
    CROSS[Cross-document<br/>Inconsistency Detector]
    THIRD[Third-party Feed Normaliser<br/>IREX · MIB · Rx · MVR → Structured Schema]
    CTX[Unified IDI Case Context<br/>All documents · All third-party data]
  end

  subgraph aop ["Agent Operating Procedure"]
    OCC[Occupation Classification<br/>Class Definitions · Duty Mix · Hazard Flags]
    FIN2[Financial Documentation<br/>Income Types · Sufficiency Rules · Cross-ref Logic]
    MED2[Medical Flagging Criteria<br/>Condition Categories · Recency · APS Adequacy]
    ESC[Escalation Criteria<br/>Thresholds · Critical Data Point Rules · Routing]
  end

  subgraph reasoning ["LLM Reasoning Layer — Governed Inference"]
    OCC_R[Occupation Analysis<br/>Title vs. Duties · Class Resolution · Duty Mix]
    FIN_R[Financial Analysis<br/>Income Verification · Sufficiency Score · Trend]
    MED_R[Medical Analysis<br/>Condition Flagging · APS Adequacy · Timeline]
    CROSS_R[Cross-evidence Synthesis<br/>Discrepancy Resolution · Requirement Generation]
    CONF[Confidence Engine<br/>Five-domain composite · IDI-calibrated weights]
    BDR[Amazon Bedrock<br/>VPC Endpoint]
    GRL[Bedrock Guardrails<br/>configured by your team]
  end

  subgraph output ["Output Layer — Your Systems"]
    REC[Structured Recommendation<br/>Decision · Exclusions · Modifications]
    SCORE[Confidence Score<br/>Composite + signal breakdown]
    REQ[Requirements List<br/>Prioritised · Document-specific · Targeted]
    AUDIT[CloudWatch Logs<br/>Step-level audit trail · Immutable]
    S3OUT[S3 Output Bucket<br/>Structured JSON]
    SOR[Policy Admin System<br/>CRM · DMS]
  end

  APP & MED & RX & FIN & MIB & MVR --> S3IN
  S3IN --> OCR --> PARSE --> NORM
  NORM --> FLAG & CROSS
  MIB & RX & MVR --> THIRD --> NORM
  FLAG & CROSS --> CTX

  OCC & FIN2 & MED2 & ESC --> OCC_R & FIN_R & MED_R

  CTX --> OCC_R
  CTX --> FIN_R
  CTX --> MED_R

  OCC_R & FIN_R & MED_R --> CROSS_R
  CROSS_R --> CONF
  OCC_R & FIN_R & MED_R & CROSS_R --> BDR
  BDR <--> GRL
  CONF --> REC & SCORE & REQ

  REC & SCORE & REQ --> S3OUT
  S3OUT & AUDIT --> SOR

  classDef boundary fill:#fafafa,stroke:#111,stroke-width:1.5px,color:#111;
  classDef store fill:#f4f4f2,stroke:#111,color:#111;
  classDef aopnode fill:#fafafa,stroke:#555,stroke-dasharray:3 3,color:#333;
  class intake,ingestion,reasoning,output boundary;
  class S3IN,S3OUT,AUDIT store;
  class OCC,FIN2,MED2,ESC aopnode;
```

*Fig. W1.1 — IDI Underwriting Agent full architecture. Every stage executes within your network boundary. The AOP governs all three reasoning domains simultaneously. Confidence scoring is the last step before output assembly — it cannot be bypassed.*

***

## 3 — Document intelligence layer

The IDI workflow processes document types and volumes that are not handled by general-purpose document processing tools. The document intelligence layer is built specifically for the IDI document ecosystem.

***

### 3.1 Multi-pass OCR pipeline

IDI application packets routinely include:

* Handwritten APS forms from treating physicians
* Faxed multi-page medical records with dot-matrix or thermal printing artefacts
* Scanned financial documents with table layouts and stamp overlays
* Occupation duty questionnaires with checkboxes, free-text fields, and signatures on the same page

The multi-pass OCR pipeline processes these document types through three sequential passes before any content is passed to the reasoning layer:

| pass                                | purpose                                                                                                                                                                                                        | output                                                                                                       |
| ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| **Pass 1 — Layout Analysis**        | Detect document structure: headers, tables, form fields, handwritten regions, and signature zones. Classify each region by content type before attempting extraction.                                          | Region map with content-type labels and geometric coordinates                                                |
| **Pass 2 — Targeted Extraction**    | Apply model-specific extraction to each region type. Printed text, handwriting, tables, and checkboxes each receive a specialized extraction model rather than a single general OCR pass.                      | Per-field extracted values with per-field OCR confidence scores                                              |
| **Pass 3 — Confidence Arbitration** | Where multiple extraction models produced competing outputs for the same field, the arbitration layer selects the highest-confidence extraction or marks the field as uncertain and flags it for human review. | Final extracted values with arbitrated confidence scores; uncertain fields marked with source ambiguity flag |

<Warning>
  A field marked as uncertain by the OCR pipeline is never silently passed to
  the reasoning layer as if it were a high-confidence extraction. The agent
  explicitly acknowledges the extraction uncertainty in its evidence citations
  and applies the appropriate confidence suppression to any finding that depends
  on that field. If the uncertain field is a material underwriting dimension —
  e.g., YTD income, elimination period, or diagnosis date — the suppression is
  sufficient to trigger a requirements flag for the specific document and field.
</Warning>

***

### 3.2 Multi-document context assembly

The IDI application packet is not a single document — it is typically seven to fifteen separate files totalling forty to one hundred pages of content. General-purpose LLM tools cannot process this volume in a single inference context, and processing documents sequentially loses the cross-document relationships that are the most valuable signal in IDI underwriting.

The context assembly layer constructs a **unified case context** that makes all documents simultaneously available to the reasoning layer:

<CardGroup cols={2}>
  <Card title="Intelligent Context Prioritisation" icon="filter">
    Rather than concatenating all documents in full, the assembly layer identifies the highest-information-density content from each document and prioritises it into the reasoning context. Full document text is available for citation; the reasoning layer operates on the structured, prioritized representation.
  </Card>

  <Card title="Entity Resolution Across Documents" icon="link">
    Named entities — the applicant, treating physicians, employers, and conditions
    — are resolved to unified identities across all documents before reasoning
    begins. This means the agent can immediately identify that "Dr. Sarah Chen" in
    the duty questionnaire and "S. Chen MD" in the APS refer to the same person,
    rather than treating them as separate entities.
  </Card>

  <Card title="Cross-Document Fact Graph" icon="diagram-project">
    Material facts extracted from multiple documents — income, occupation,
    conditions, dates, benefit amounts — are assembled into a structured fact
    graph that explicitly tracks which documents support, contradict, or are
    silent on each fact. The reasoning layer operates on this graph, not on raw
    document text.
  </Card>

  <Card title="Third-Party Feed Integration" icon="plug">
    Structured feeds from IREX, MIB, pharmacy claims databases, and motor vehicle records are normalized into the same unified schema as the submitted documents. The agent treats third-party data and applicant-submitted documents as two independent evidence streams and explicitly identifies any divergence between them.
  </Card>
</CardGroup>

***

### 3.3 Third-party data normalisation

Third-party data sources for IDI cases arrive in a variety of formats — IREX returns proprietary structured XML, MIB returns coded records requiring interpretation against the MIB code glossary, pharmacy claims databases return structured but insurer-specific schemas, and MVR returns vary by state and provider.

The normalisation layer converts all of these feeds into the unified IDI case schema before any reasoning step:

```json theme={null}
{
  "third_party_data": {
    "irex": {
      "attending_physician_statements": [
        {
          "physician_name": "Dr. Sarah Chen, MD",
          "specialty": "Internal Medicine",
          "last_exam_date": "2024-11-03",
          "conditions_disclosed": ["hypertension", "hyperlipidemia"],
          "medications_prescribed": ["lisinopril_10mg", "atorvastatin_20mg"],
          "extraction_confidence": 0.97
        }
      ]
    },
    "mib": {
      "codes_found": ["A06", "D14"],
      "codes_interpreted": [
        {
          "code": "A06",
          "category": "cardiovascular",
          "significance": "moderate"
        },
        { "code": "D14", "category": "musculoskeletal", "significance": "low" }
      ],
      "cross_reference_status": "discrepancy_detected",
      "discrepancy_detail": "MIB code A06 not corroborated by applicant-submitted APS"
    },
    "rx_history": {
      "medications": [
        {
          "drug_name": "sertraline",
          "drug_class": "SSRI",
          "fill_dates": ["2023-04-12", "2023-07-08", "2023-10-15"],
          "prescribing_physician": "Dr. James Park, MD",
          "significance_flag": "mental_health_indicator",
          "disclosed_on_application": false
        }
      ]
    }
  }
}
```

<Note>
  The cross-reference status field in the normalized third-party schema is
  computed by the ingestion layer, not the reasoning layer. This means the agent
  enters the reasoning phase already aware of structural discrepancies between
  third-party data and submitted documents — it does not discover them
  mid-reasoning. This architecture prevents the reasoning layer from anchoring
  on submitted documents before evaluating third-party data.
</Note>

***

## 4 — Confidence engine calibration for IDI

The confidence engine architecture is described in [the Confidence Engine reference](/agents/confidence/engine). This section covers the IDI-specific elements: how the five signal domains are weighted for this workflow, and how the engine is validated against an IDI-specific synthetic dataset before any model version is certified for production.

***

### 4.1 IDI signal domain weights

The five confidence signal domains carry different weights in the IDI composite than they would in a different workflow agent. The weight set is calibrated to reflect the relative evidential importance of each domain in IDI underwriting specifically.

| signal domain                   | IDI weight | rationale                                                                                                                                                                                                                                                                          |
| ------------------------------- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Cross-Evidence Consistency**  | 30%        | Income, occupation, and medical history are stated on the application and independently evidenced by multiple documents and third-party feeds. Discrepancies between stated and evidenced facts are the highest-value underwriting signal in IDI — the weight reflects this.       |
| **AOP Coverage**                | 25%        | IDI cases regularly surface occupation and medical combinations that lie at the edge of AOP coverage. AOP gaps in IDI are more likely to be genuinely ambiguous than in other workflows and carry a higher escalation burden.                                                      |
| **Document Extraction Quality** | 20%        | IDI packets include some of the most challenging document types in insurance: handwritten APS forms, faxed records, and multi-page financials with complex layouts. Extraction quality directly determines whether the material facts are recoverable at all.                      |
| **Inference Quality**           | 15%        | The IDI reasoning chain is longer and more multi-step than simpler workflows. Model uncertainty accumulates across occupation, financial, and medical reasoning passes. This domain captures that accumulated uncertainty.                                                         |
| **External Data Alignment**     | 10%        | Third-party data alignment is a critical signal but is treated as a detractor amplifier rather than a standalone domain — a MIB discrepancy alone does not resolve the case, but significantly amplifies the weight of any corroborating inconsistency in the submitted documents. |

***

### 4.2 IDI-specific detractors

In addition to the standard detractor taxonomy, the IDI workflow includes several detractors that are specific to this product line and do not appear in other agent configurations.

<CardGroup cols={2}>
  <Card title="Undisclosed Rx — Mental Health Indicator" icon="capsules">
    A prescription in the normalized Rx history falls within a drug class associated with mental health treatment (SSRIs, SNRIs, atypical antipsychotics, mood stabilisers) and was not disclosed on the application form. This detractor fires regardless of whether the condition itself is ultimately material — non-disclosure of any medication is an independent underwriting finding.
  </Card>

  <Card title="Occupation Class Dispute" icon="briefcase">
    The agent's resolved occupation class differs from the occupation class
    submitted on the application. This detractor carries a high severity weight
    because benefit amounts and premium rates are directly tied to occupation
    class — an incorrect class on the submitted application creates financial
    exposure independent of any other finding.
  </Card>

  <Card title="Income Trend Adverse" icon="chart-line-down">
    The income trend across the available documentation years is materially
    negative — defined in the AOP as a year-over-year decline exceeding the
    configured adverse trend threshold. Declining income trajectories affect both
    benefit amount calculation and the sustainability of the proposed premium
    obligation and are surfaced as a financial analysis finding.
  </Card>

  <Card title="APS Specialty Mismatch" icon="user-doctor">
    The Attending Physician Statement was authored by a physician whose stated
    specialty is not consistent with the condition being reported. For example, an
    APS reporting on a complex psychiatric history authored by a general
    practitioner, where the AOP requires a specialist APS for that condition
    category, triggers this detractor and generates a targeted specialist APS
    requirement.
  </Card>

  <Card title="Reinsurance Treaty Boundary" icon="scale-balanced">
    The case characteristics — benefit amount, occupation class, age, benefit
    period combination — place it at or near a boundary defined in the applicable
    reinsurance treaty. Cases within the AOP-configured margin of a treaty
    boundary receive a confidence suppression and are flagged for reinsurer
    consultation before disposition.
  </Card>

  <Card title="Inconsistent Activity Level" icon="activity">
    The applicant's disclosed occupation (e.g., sedentary office role, Class 4A) is inconsistent with activities or physical findings described in the APS or implied by MVR data. This detractor is specific to the intersection of medical and occupational evidence and does not fire from either source independently.
  </Card>
</CardGroup>

***

## 5 — IDI benchmark results

The agent is evaluated against an IDI-specific synthetic validation dataset before any model version is certified for production. This section describes the validation methodology, the dataset, and the certified benchmark results for the current production model.

***

### 5.1 Validation methodology

The IDI synthetic validation dataset is constructed by Layerup's underwriting implementation team to represent the full distribution of case complexity encountered in production IDI underwriting. Cases in the dataset are assigned ground-truth labels by experienced IDI underwriters — not derived from model outputs.

The dataset is structured to deliberately stress every agent capability:

| dataset segment                      | construction method                                                                                               | purpose                                                                                             |
| ------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| **Clean cases**                      | Synthetically generated complete packets with no deliberate defects                                               | Establish baseline precision; agent should resolve these at high confidence                         |
| **OCR-degraded cases**               | Clean cases with intentionally degraded document quality (rotation, noise, reduced DPI, handwriting substitution) | Validate multi-pass OCR pipeline under realistic document quality conditions                        |
| **Cross-document discrepancy cases** | Cases with deliberate income, occupation, or medical inconsistencies inserted across document pairs               | Validate cross-evidence consistency detection; agent must surface all inserted discrepancies        |
| **AOP gap cases**                    | Cases with occupation or medical combinations not covered by the AOP                                              | Validate that the agent surfaces gaps rather than extrapolating; no recommendation should be issued |
| **Undisclosed condition cases**      | Cases where Rx history or MIB data reveals conditions not on the application form                                 | Validate third-party data normalisation and non-disclosure detection                                |
| **Reinsurance boundary cases**       | Cases synthetically placed at treaty boundary conditions                                                          | Validate boundary detection and appropriate confidence suppression                                  |

Before any AOP version or model version is promoted to production, the agent must pass all six benchmark checks listed in Section 5.2. Failure on any single check blocks promotion.

***

### 5.2 Certified benchmark results — current production model

The following benchmarks are the certified results for the current production model version against the IDI synthetic validation dataset. Results are expressed as the mean and interquartile range (IQR) across three independent validation runs.

**Recommendation accuracy** — overall alignment between agent recommendation and ground-truth underwriter disposition:

| decision type                     | precision | recall    | F1        |
| --------------------------------- | --------- | --------- | --------- |
| Approve as Applied                | 99.1%     | 99.3%     | 99.2%     |
| Approve with Exclusion / Modified | 98.4%     | 98.7%     | 98.5%     |
| Defer Pending Requirements        | 98.8%     | 99.0%     | 98.9%     |
| Escalate                          | 99.6%     | 99.4%     | 99.5%     |
| Decline                           | 99.2%     | 99.1%     | 99.1%     |
| **Overall weighted**              | **99.0%** | **99.1%** | **99.0%** |

**Cross-document discrepancy detection** — on the discrepancy segment of the validation set:

| discrepancy type                                 | detection rate | false positive rate |
| ------------------------------------------------ | -------------- | ------------------- |
| Income discrepancy (cross-document)              | 99.4%          | 0.5%                |
| Occupation class dispute                         | 99.1%          | 0.4%                |
| Medical timeline inconsistency                   | 98.3%          | 0.8%                |
| Undisclosed Rx — application vs. pharmacy record | 99.7%          | 0.2%                |
| MIB code — undisclosed condition indicator       | 98.9%          | 0.6%                |
| APS specialty mismatch                           | 98.6%          | 0.5%                |

**Confidence score calibration** — Brier score and Expected Calibration Error (ECE) measure how accurately the composite confidence score predicts actual recommendation accuracy:

| metric                                 | IDI result | interpretation                                                                                                                                                    |
| -------------------------------------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Brier Score                            | 0.009      | Lower is better; 0 = perfect calibration. A score of 0.009 indicates the confidence score is a highly reliable predictor of recommendation accuracy.              |
| Expected Calibration Error (ECE)       | 0.8%       | The average gap between the stated confidence and the empirical accuracy across the confidence range.                                                             |
| Cases correctly escalated at threshold | 99.6%      | Proportion of validation cases that should have received human review (ground truth) that did receive hard escalation from the confidence engine.                 |
| Cases incorrectly auto-resolved        | 0.1%       | Proportion of validation cases where the agent produced a recommendation at or above the high threshold but the ground-truth disposition was Escalate or Decline. |

**OCR extraction accuracy** — on the OCR-degraded segment of the validation set:

| document type                       | field-level extraction accuracy | extraction floor hit rate |
| ----------------------------------- | ------------------------------- | ------------------------- |
| Handwritten APS                     | 98.1%                           | 1.4%                      |
| Faxed medical records               | 98.7%                           | 0.9%                      |
| Scanned financials (complex layout) | 99.2%                           | 0.6%                      |
| Duty questionnaires (mixed format)  | 98.4%                           | 1.1%                      |
| Standard printed PDFs               | 99.8%                           | 0.1%                      |

<Note>
  "Extraction floor hit rate" is the proportion of pages where OCR confidence
  falls below the configured extraction floor threshold, triggering an
  uncertainty flag rather than passing the extracted value to the reasoning
  layer. A higher floor hit rate on handwritten documents reflects a
  conservative extraction policy — the agent prefers to flag and require human
  verification over silently passing low-confidence extractions.
</Note>

**AOP gap handling** — on the AOP gap segment of the validation set:

| gap scenario                                | correct gap surface rate | hallucination rate |
| ------------------------------------------- | ------------------------ | ------------------ |
| Occupation outside AOP class coverage       | 100%                     | 0%                 |
| Medical condition with no AOP flagging rule | 99.8%                    | 0%                 |
| Multi-jurisdiction regulatory conflict      | 99.1%                    | 0%                 |
| Intra-AOP rule conflict                     | 99.4%                    | 0%                 |

<Warning>
  A hallucination rate of 0% in the AOP gap segment means the agent produced
  zero instances where it issued a recommendation by extrapolating beyond its
  AOP coverage. This is a hard architectural invariant: the agent is not
  permitted to issue a recommendation for a case dimension where no AOP rule
  applies. The zero result is expected and is enforced by the output assembly
  layer — it validates AOP coverage before assembling any recommendation output.
</Warning>

***

### 5.3 Score distribution under production conditions

The certified baseline confidence score distribution across the validation set establishes the reference against which all production AOP updates are evaluated. Any AOP promotion that shifts the distribution beyond the configured tolerance band is blocked by the CI/CD test harness.

| confidence tier                               | validation set distribution | production target range |
| --------------------------------------------- | --------------------------- | ----------------------- |
| Auto-Resolve (≥ high threshold)               | 58.4%                       | 55–65%                  |
| Soft Review (mid threshold to high threshold) | 28.7%                       | 25–35%                  |
| Hard Escalation (\< low threshold)            | 12.9%                       | 10–18%                  |

<Note>
  The production target range is intentionally wide enough to accommodate
  genuine variation across customer AOP configurations while being narrow enough
  to detect implausible distributions. A distribution concentrated above 70%
  auto-resolve would suggest a permissive AOP configuration that is not applying
  appropriate sensitivity to IDI-specific risk signals. A distribution with over
  25% hard escalations would suggest an AOP that is not covering the case
  population adequately.
</Note>

***

## 6 — Structured output schema

Every IDI case processed by the agent produces a single structured JSON output payload. The following describes the top-level schema and the IDI-specific fields.

```json theme={null}
{
  "case_id": "IDI-2024-110482",
  "policy_number": "DI-001-482910",
  "processed_at": "2024-12-04T14:37:22Z",
  "processing_time_seconds": 847,

  "applicant": {
    "name": "Jane Applicant",
    "age": 41,
    "gender": "Female",
    "residence_state": "NY",
    "citizenship": "US Citizen"
  },

  "occupation_analysis": {
    "submitted_job_title": "Emergency Physician",
    "submitted_occupation_class": "4A",
    "resolved_occupation_class": "4A",
    "class_dispute": false,
    "duty_mix": {
      "manual_percent": 62,
      "supervisory_percent": 18,
      "clerical_percent": 20,
      "resolution_confidence": 0.91,
      "sources": ["duty_questionnaire.pdf", "aps_dr_chen.pdf"]
    },
    "hazard_flags": [],
    "title_vs_duties_mismatch": false
  },

  "financial_analysis": {
    "income_type": "W-2 employment",
    "prior_year_income": 310000,
    "two_year_prior_income": 295000,
    "income_trend": "stable",
    "income_sufficiency_score": 0.96,
    "monthly_benefit_supportable": 15000,
    "financial_requirements": []
  },

  "medical_analysis": {
    "conditions_identified": [
      {
        "condition": "essential hypertension",
        "category": "cardiovascular",
        "source_documents": ["aps_dr_chen.pdf"],
        "last_treatment_date": "2024-10-15",
        "significance": "low",
        "exclusion_recommendation": null
      }
    ],
    "aps_adequacy": "adequate",
    "undisclosed_rx_flags": [],
    "mib_discrepancies": [],
    "timeline_inconsistencies": []
  },

  "ai_recommendation": {
    "decision": "Approve as Applied",
    "confidence_score": 94,
    "exclusions_recommended": [],
    "modifications_recommended": [],
    "requirements": [],
    "escalation_required": false,
    "confidence_signal_breakdown": {
      "composite_score": 94,
      "signals": [
        {
          "domain": "cross_evidence_consistency",
          "domain_score": 96,
          "weight": 0.3
        },
        { "domain": "aop_coverage", "domain_score": 100, "weight": 0.25 },
        {
          "domain": "document_extraction_quality",
          "domain_score": 93,
          "weight": 0.2
        },
        { "domain": "inference_quality", "domain_score": 91, "weight": 0.15 },
        {
          "domain": "external_data_alignment",
          "domain_score": 88,
          "weight": 0.1
        }
      ]
    }
  },

  "evidence_citations": [
    {
      "finding": "Occupation class resolved to 4A",
      "sources": [
        {
          "document": "application_form.pdf",
          "page": 2,
          "field": "occupation_class"
        },
        {
          "document": "duty_questionnaire.pdf",
          "page": 1,
          "field": "duties_narrative"
        }
      ]
    }
  ],

  "documents_processed": 8,
  "total_pages_reviewed": 61,
  "aop_version": "v2.4.1",
  "model_version": "layerup-idi-v3"
}
```

***

## 7 — KPI impact model

The IDI Underwriting Agent is designed to produce measurable, auditable improvements across the underwriting operations KPIs that matter most. The following describes the KPI model and the mechanisms by which each improvement is achieved.

***

### 7.1 Throughput & cycle time

<CardGroup cols={2}>
  <Card title="Case Throughput" icon="gauge">
    The agent runs in parallel across all queued cases simultaneously — there is no single-threaded bottleneck. Throughput scales with compute configuration, not with headcount. New business volumes that historically required hiring or overtime can be absorbed by scaling the agent's task configuration.
  </Card>

  <Card title="Cycle Time Compression" icon="clock">
    The agent completes the initial case review — including document extraction,
    cross-referencing, medical flagging, financial analysis, and requirement
    generation — in minutes rather than days. The human underwriter receives a
    fully prepared case record rather than a raw document stack, compressing the
    total elapsed time from submission to decision.
  </Card>

  <Card title="Requirement Latency Reduction" icon="list-check">
    The requirements list is generated at the end of the agent's first pass —
    before any human underwriter touches the case. Outstanding requirements reach
    the submitting agent or applicant earlier, reducing the time cases spend in
    pending status due to incomplete documentation.
  </Card>

  <Card title="In-Good-Order Rate Improvement" icon="circle-check">
    By generating targeted, document-specific requirements rather than generic follow-up requests, the agent reduces the round-trip count needed to resolve requirements. Agents and applicants receive precise instructions — "provide APS from treating psychiatrist Dr. Park covering the period 2022–2024" rather than "provide psychiatric APS."
  </Card>
</CardGroup>

***

### 7.2 Underwriting quality & consistency

<CardGroup cols={2}>
  <Card title="Consistent AOP Application" icon="equals">
    The agent applies the AOP with zero variance across every case. No case is processed differently because of the time of day, the underwriter's experience level, or workload pressure. Occupation class decisions, financial thresholds, and medical flagging criteria are applied identically on every case.
  </Card>

  <Card title="Non-Disclosure Detection Rate" icon="magnifying-glass">
    The cross-reference between applicant-submitted documents and third-party data
    feeds (Rx history, MIB) surfaces non-disclosures that human review of
    submitted documents alone cannot identify. The agent's validated
    non-disclosure detection rate exceeds 98% on Rx-sourced indicators in the
    benchmark dataset.
  </Card>

  <Card title="Intra-AOP Conflict Surface" icon="arrows-split-up-and-left">
    As the agent processes cases against the AOP, it surfaces situations where AOP
    clauses conflict — inconsistencies that may have existed in the underlying SOP
    for years without being identified because individual underwriters resolved
    them informally. These surfaces provide a systematic method for identifying
    and resolving SOP ambiguities before they affect decisions.
  </Card>

  <Card title="Decision Auditability" icon="receipt">
    Every recommendation the agent issues includes a full evidence citation chain from raw document to extracted fact to reasoning step to decision. Audit and compliance review of any historical case can trace the complete reasoning without re-running the agent or relying on underwriter notes.
  </Card>
</CardGroup>

***

### 7.3 Human underwriter productivity

The agent is designed to make the underwriter's remaining work higher-value, not to create additional review burdens. Cases that reach a human underwriter from the agent queue arrive with:

* A structured recommendation with occupation class resolution, income verification, and medical flagging already completed
* A prioritized requirements list with document-specific, targeted requirements
* A confidence signal breakdown that tells the underwriter exactly which evidence domain is weakest and why
* Full evidence citations for every finding, so the underwriter can immediately navigate to the relevant document and page

The net effect is that the underwriter's cognitive load on a case prepared by the agent is concentrated on the genuinely ambiguous elements — the ones the confidence engine surfaced as requiring expert judgment — rather than spread across the full document stack.

| underwriter activity               | before agent               | with agent                                                                          |
| ---------------------------------- | -------------------------- | ----------------------------------------------------------------------------------- |
| Document review (full stack)       | 45–90 min per case         | 0 — agent produces indexed, cited findings                                          |
| Occupation class determination     | 15–30 min per complex case | 0 for resolved cases; 10 min review for flagged cases                               |
| Financial documentation review     | 20–40 min per case         | 0 for complete cases; focused review of flagged discrepancies only                  |
| Requirement letter drafting        | 15–25 min per case         | 0 — agent generates targeted requirement list for underwriter approval              |
| Cross-referencing third-party data | 20–35 min per case         | 0 — agent normalizes and cross-references all feeds before case reaches underwriter |
| **Focused case judgment**          | Residual after above       | **Primary underwriter activity — exercised on the cases the agent cannot resolve**  |

***

### 7.4 Risk & compliance KPIs

<CardGroup cols={2}>
  <Card title="False Approval Rate" icon="shield-x">
    The confidence engine's calibrated escalation logic ensures that cases the agent is not entitled to approve do not auto-resolve. The validated false approval rate in the IDI benchmark set — cases where the agent issued an Approve recommendation but the ground-truth disposition was Escalate or Decline — is 0.4%.
  </Card>

  <Card title="Escalation Capture Rate" icon="arrow-up-right">
    98.7% of validation cases with a ground-truth escalation disposition received
    a hard escalation from the agent. Escalations are never suppressed by the
    output assembly layer — the escalation flag is written before the
    recommendation payload is assembled.
  </Card>

  <Card title="Regulatory Exam Readiness" icon="file-check">
    The structured, evidence-cited output payload and the immutable CloudWatch
    audit trail give your compliance team a complete, immediately accessible
    record for any case under regulatory examination — without requiring
    underwriter recall or document reconstruction.
  </Card>

  <Card title="SOP Enforcement Consistency" icon="book-open-check">
    The agent applies your AOP-encoded SOP identically across 100% of cases. Enforcement consistency across product lines, distribution channels, and states is a direct output of AOP-governed operation — not a function of underwriter training or manager oversight.
  </Card>
</CardGroup>

***

## 8 — Rollout & human-in-the-loop model

The IDI Underwriting Agent is deployed through a structured rollout model that preserves human decision authority at every stage. The agent's role expands progressively as confidence in its calibration is established through real production data.

```mermaid theme={null}
flowchart LR
  P1["Phase 1 — Shadow Mode<br/>Agent runs on every case<br/>Underwriter makes all decisions<br/>Agent output visible but advisory only"]
  P2["Phase 2 — Assist Mode<br/>Agent output is primary review surface<br/>Underwriter reviews all recommendations<br/>Accepts, modifies, or overrides each"]
  P3["Phase 3 — Auto-Resolve (High Confidence)<br/>Cases above high confidence threshold<br/>auto-resolve without mandatory review<br/>Underwriter reviews queue of flagged cases"]
  P4["Phase 4 — Full Deployment<br/>Auto-resolve + soft review queue<br/>Hard escalations route to senior underwriter<br/>Ongoing benchmark monitoring active"]

  P1 -->|"Calibration validated<br/>against shadow outcomes"| P2
  P2 -->|"Recommendation alignment<br/>confirmed > threshold"| P3
  P3 -->|"Auto-resolve error rate<br/>confirmed < tolerance"| P4

  classDef phase fill:#fafafa,stroke:#111,color:#111;
  class P1,P2,P3,P4 phase;
```

*Fig. W1.2 — IDI rollout phases. Each phase transition is gated by a measurable calibration check. No phase transition is automatic — your underwriting governance team approves each gate.*

<Note>
  Human decision authority is preserved unconditionally throughout Phases 1 and
  2, and for all hard-escalation cases in Phases 3 and 4. The agent does not
  make a binding underwriting decision at any phase — it produces a
  recommendation. Your underwriting governance team configures the confidence
  threshold at which recommendations are treated as actionable without mandatory
  individual review.
</Note>
