> ## Documentation Index
> Fetch the complete documentation index at: https://docs.uselayerup.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Mortgage Insurance Underwriting Agent

> End-to-end architecture, document intelligence pipeline, confidence calibration, and performance benchmarks for the mortgage insurance underwriting workflow.

# Mortgage Insurance Underwriting Agent — architecture, benchmarks & workflow deep dive.

The Mortgage Insurance (MI) Underwriting Agent is a fully sovereign, AOP-governed reasoning workload purpose-built for new business underwriting across private mortgage insurance product lines — including flow, bulk, delegated, and non-delegated submissions. It ingests the complete loan file — Uniform Residential Loan Application (URLA / 1003), credit report, income documentation, bank statements, appraisal, automated underwriting system findings, and guideline overlays — processes every document simultaneously in a unified reasoning context, and produces a structured, evidence-cited, confidence-scored recommendation within a single automated pass.

This page covers the MI-specific architecture, the insurance ontology embedded in the Agent Operating Procedure, how the confidence engine is calibrated and validated for this workflow, independently measured benchmarks against the MI-specific synthetic validation set, and the KPI impact model for underwriting operations.

Mortgage insurance underwriting is distinguished from other personal-lines underwriting workflows by its income-document variety (W-2, pay stubs, bank statements, tax returns, VOE), its credit-report density, its LTV / DTI / FICO eligibility math against published guidelines and overlays, and its delegated versus non-delegated referral discipline. The AOP schema encodes all of these dimensions as first-class underwriting objects rather than general document extraction targets.

| attribute          | value                                                                                                                                                                                            |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| workflow           | Mortgage Insurance — New Business Underwriting                                                                                                                                                   |
| case types         | Flow · Bulk · Delegated · Non-delegated · Rate-quote / pre-qualification                                                                                                                         |
| document intake    | URLA / 1003 · 1008 · Credit report (tri-merge) · W-2 · Pay stubs · Bank statements · Tax returns · VOE · Appraisal · AUS findings (DU / LPA) · Title / HOI · Flood cert · Letters of explanation |
| decision outputs   | Approve as Applied · Approve with Conditions / Overlay · Refer Non-Delegated · Defer Pending Requirements · Escalate · Decline / Ineligible                                                      |
| confidence scoring | Composite 0–100 · Five independent signal domains · MI-calibrated weight set                                                                                                                     |
| deployment model   | Private VPC / VNet — your cloud account, your network                                                                                                                                            |

***

## 1 — MI-specific Agent Operating Procedure

The AOP is the agent's underwriting contract. For the MI workflow, it is not a generic document-processing instruction set — it is built on an **insurance-native ontology** that Layerup has assembled from first-principles mortgage insurance underwriting practice, including income documentation hierarchies, credit and capacity rules, LTV / DTI / FICO eligibility tables, occupancy and property-type overlays, and delegated authority boundaries. Your enterprise SOP and published guideline overlays are layered on top of this baseline, not substituted for it.

The MI AOP schema covers three primary domains that are specific to mortgage insurance underwriting and do not exist in Layerup's other workflow agents.

***

### 1.1 Income and employment verification

MI eligibility and pricing are tied to documented, stable income. The AOP encodes income verification requirements by income type, documentation hierarchy, and cross-reference logic.

| income type                   | primary document                          | secondary document             | cross-reference logic                                                                       |
| ----------------------------- | ----------------------------------------- | ------------------------------ | ------------------------------------------------------------------------------------------- |
| W-2 wage / salary             | Most recent two years W-2                 | YTD pay stubs (30-day recency) | YTD pay stub annualised and reconciled against W-2; variance above AOP tolerance flagged    |
| Hourly / variable hours       | Two years W-2 + YTD pay stubs             | VOE (hours and YTD)            | Two-year average applied; declining hours trend triggers additional VOE                     |
| Bonus / overtime / commission | Two-year history on W-2 / pay stubs       | Employer letter or VOE         | Included at AOP-defined inclusion percentage; declining trend excluded or haircut applied   |
| Self-employed / Schedule C    | Two years federal returns (all schedules) | YTD P\&L or CPA letter         | Year-over-year variance flagged if > AOP-defined tolerance; add-backs applied per guideline |
| Partnership / S-Corp / K-1    | K-1 for two years + business returns      | CPA letter                     | Active participation verification; passive K-1 income excluded                              |
| Rental / other                | Schedule E + lease or tax returns         | Bank-statement deposit support | Vacancy / expense haircut per AOP; undocumented deposits excluded from qualifying income    |
| Asset-based / asset depletion | Bank / brokerage statements (seasoned)    | Account ownership evidence     | Seasoning and large-deposit sourcing applied; ineligible assets excluded                    |

The agent does not compute a binary pass/fail against these requirements. It produces an **income sufficiency score** for each applicable income type, identifies which documents are present and which are absent, and generates a prioritized requirements list ranked by the evidentiary gap's impact on the confidence composite.

<CardGroup cols={2}>
  <Card title="Pay Stub Recency & Continuity" icon="calendar">
    The AOP encodes the recency window and consecutive-period requirements for pay stubs. Gaps between pay periods, a stub older than the configured recency floor, or a YTD figure that cannot be reconciled to the stated base pay are surfaced as first-class income findings — not buried in a generic document-quality flag.
  </Card>

  <Card title="Bank-Statement Deposit Analysis" icon="building-columns">
    For bank-statement income or large-deposit sourcing, the agent extracts
    transaction tables, identifies recurring deposits, flags NSF / overdraft
    patterns, and sources large deposits against the loan file. Unsourced large
    deposits and income that appears only as cash deposits without a matching VOE
    or tax return are surfaced as capacity findings.
  </Card>

  <Card title="W-2 vs. Pay Stub Reconciliation" icon="arrows-left-right">
    W-2 Box 1 / Box 5 figures are reconciled against annualised YTD pay-stub
    earnings and the income stated on the URLA. Material mismatches — different
    employers, a mid-year job change without a VOE, or a stated income that the
    documents do not support — fire as income-domain discrepancies with both
    source documents cited.
  </Card>

  <Card title="VOE / The Work Number Alignment" icon="briefcase">
    When a third-party employment verification feed is present, the agent treats it as an independent evidence stream against the borrower-submitted W-2, pay stubs, and URLA employer fields. Hire date, job title, and YTD earnings divergences are surfaced before the reasoning layer anchors on the submitted file.
  </Card>
</CardGroup>

***

### 1.2 Credit and capacity

Credit-report density and DTI construction are the second primary risk driver. The AOP encodes how tradelines, inquiries, public records, and stated liabilities are assembled into a capacity picture.

<CardGroup cols={2}>
  <Card title="Tri-Merge Credit Normalisation" icon="file-invoice">
    Equifax, Experian, and TransUnion tradelines are resolved to a single liability set. Duplicate tradelines, authorized-user accounts, and medical collections are classified per the AOP inclusion rules. The representative credit score is selected by the AOP-configured method (middle of three, lower of two) — not by a generic "best score" heuristic.
  </Card>

  <Card title="DTI Construction" icon="chart-pie">
    Housing expense (proposed PITIA) and monthly liabilities are assembled from
    the credit report, URLA, and supporting documents. The agent distinguishes
    qualifying DTI from stated DTI, applies AOP rules for installment accounts
    with fewer than N remaining payments, and flags omitted liabilities that
    appear on credit but not on the 1003.
  </Card>

  <Card title="Credit Event Recency" icon="clock-rotate-left">
    Bankruptcy, foreclosure, short sale, deed-in-lieu, and significant derogatory
    events are dated and evaluated against AOP seasoning windows. An event inside
    the window is a hard eligibility finding; an event at the window boundary
    receives a confidence suppression and a referral flag.
  </Card>

  <Card title="Inquiry & Undisclosed Debt" icon="magnifying-glass">
    Recent inquiries without a corresponding new tradeline, and new tradelines
    opened after the credit-report date that appear on a refresh, are surfaced as
    undisclosed-debt risk. The agent does not assume an inquiry is benign.
  </Card>

  <Card title="Letter of Explanation Adequacy" icon="file-lines">
    Borrower letters of explanation — often handwritten or wet-ink — are evaluated
    for whether they address the specific credit finding cited (late pays, large
    deposits, employment gap). A generic letter that does not address the flagged
    item generates a targeted follow-up requirement.
  </Card>

  <Card title="Occupancy vs. Credit Footprint" icon="house">
    The declared occupancy (primary, second home, investment) is cross-checked against address history on the credit report, URLA residences, and appraisal occupancy indicators. Occupancy inconsistencies are first-class eligibility findings.
  </Card>
</CardGroup>

***

### 1.3 Property, collateral, and guideline eligibility

MI coverage is collateral-backed. The AOP encodes LTV / CLTV construction, property-type and occupancy eligibility, appraisal sufficiency, and the delegated versus non-delegated authority boundary.

<CardGroup cols={2}>
  <Card title="LTV / CLTV / HCLTV Construction" icon="percent">
    Base LTV is computed from the lesser of purchase price or appraised value against the insured loan amount, with subordinate financing included per the AOP. Rounding conventions, financed MI, and closing-cost credits that affect the insurable balance are applied as encoded — not inferred.
  </Card>

  <Card title="Guideline & Overlay Screening" icon="table">
    FICO floor, max DTI, max LTV by occupancy and property type, reserve
    requirements, and product overlays are evaluated as a structured eligibility
    matrix. A file that passes the published guideline but fails a company overlay
    is classified as overlay-ineligible, not as a generic decline.
  </Card>

  <Card title="Appraisal Sufficiency" icon="image">
    The agent extracts subject property data, comparable grid, condition/quality
    ratings, and appraisal conditions. A value that does not support the requested
    LTV, a condition rating below the AOP floor, or a comparable set that fails
    the AOP distance / recency rules generates a collateral finding and, where
    configured, an appraisal review requirement.
  </Card>

  <Card title="AUS Findings Alignment" icon="robot">
    DU / LPA findings are treated as an independent evidence stream. An
    Approve/Eligible AUS that the document file does not support — or a
    Refer/Caution AUS on a file the documents would otherwise clear — is surfaced
    as an AUS-file discrepancy, not silently overridden.
  </Card>

  <Card title="Delegated Authority Boundary" icon="scale-balanced">
    Delegated submissions are evaluated against the delegated authority matrix
    (FICO, LTV, DTI, occupancy, property type, and exception count). Files at or
    beyond the boundary are routed as Refer Non-Delegated with the specific matrix
    cell cited. The agent is not permitted to auto-approve through a
    delegated-authority wall.
  </Card>

  <Card title="Insurance & Title Conditions" icon="file-contract">
    HOI coverage amount and named insured, flood-zone determination, and title exceptions that affect insurability are checked against AOP minimums. Missing HOI, inadequate dwelling coverage relative to the loan amount, or a flood zone requiring coverage without a matching policy are requirement-generating findings.
  </Card>
</CardGroup>

***

## 2 — MI workflow architecture

The following diagram describes the complete data and reasoning flow for a single MI underwriting case. Every component executes within your cloud account boundary.

```mermaid theme={null}
flowchart TD
  subgraph intake ["Intake Layer — Your Systems"]
    APP[URLA / 1003<br/>+ 1008 / Transmittal]
    CR[Credit Report<br/>Tri-merge]
    INC[Income Documents<br/>W-2 · Pay Stubs · Tax Returns]
    BANK[Bank Statements<br/>Asset / Income]
    APR[Appraisal<br/>+ Title · HOI · Flood]
    AUS[AUS Findings<br/>DU / LPA]
    S3IN[S3 Input Bucket<br/>CMK-encrypted]
  end

  subgraph ingestion ["Document Intelligence Layer — Private Subnet"]
    OCR[Multi-pass OCR Pipeline<br/>Bank Statements · Pay Stubs · Credit · Handwriting]
    PARSE[Document Parser<br/>Layout · Tables · Handwriting · Signatures]
    NORM[Schema Normalisation<br/>MI Loan-File Schema]
    FLAG[Ambiguity Flagger<br/>Low-confidence fields marked]
    CROSS[Cross-document<br/>Inconsistency Detector]
    THIRD[Third-party Feed Normaliser<br/>Credit · AUS · VOE · Flood → Structured Schema]
    CTX[Unified MI Case Context<br/>All documents · All third-party data]
  end

  subgraph aop ["Agent Operating Procedure"]
    INC2[Income Verification<br/>Doc Hierarchy · Reconciliation · Seasoning]
    CRED[Credit & Capacity<br/>Tradelines · DTI · Credit Events]
    ELIG[Eligibility & Collateral<br/>LTV · Overlays · Delegated Boundary]
    ESC[Escalation Criteria<br/>Thresholds · Authority Rules · Routing]
  end

  subgraph reasoning ["LLM Reasoning Layer — Governed Inference"]
    INC_R[Income Analysis<br/>Sufficiency · Trend · Reconciliation]
    CRED_R[Credit Analysis<br/>DTI · Derogatory · Undisclosed Debt]
    ELIG_R[Eligibility Analysis<br/>LTV · Overlay · Appraisal · AUS]
    CROSS_R[Cross-evidence Synthesis<br/>Discrepancy Resolution · Requirement Generation]
    CONF[Confidence Engine<br/>Five-domain composite · MI-calibrated weights]
    BDR[Amazon Bedrock<br/>VPC Endpoint]
    GRL[Bedrock Guardrails<br/>configured by your team]
  end

  subgraph output ["Output Layer — Your Systems"]
    REC[Structured Recommendation<br/>Decision · Conditions · Overlay]
    SCORE[Confidence Score<br/>Composite + signal breakdown]
    REQ[Requirements List<br/>Prioritised · Document-specific · Targeted]
    AUDIT[CloudWatch Logs<br/>Step-level audit trail · Immutable]
    S3OUT[S3 Output Bucket<br/>Structured JSON]
    SOR[MI Underwriting Workbench<br/>LOS · Pricing]
  end

  APP & CR & INC & BANK & APR & AUS --> S3IN
  S3IN --> OCR --> PARSE --> NORM
  NORM --> FLAG & CROSS
  CR & AUS --> THIRD --> NORM
  FLAG & CROSS --> CTX

  INC2 & CRED & ELIG & ESC --> INC_R & CRED_R & ELIG_R

  CTX --> INC_R
  CTX --> CRED_R
  CTX --> ELIG_R

  INC_R & CRED_R & ELIG_R --> CROSS_R
  CROSS_R --> CONF
  INC_R & CRED_R & ELIG_R & CROSS_R --> BDR
  BDR <--> GRL
  CONF --> REC & SCORE & REQ

  REC & SCORE & REQ --> S3OUT
  S3OUT & AUDIT --> SOR

  classDef boundary fill:#fafafa,stroke:#111,stroke-width:1.5px,color:#111;
  classDef store fill:#f4f4f2,stroke:#111,color:#111;
  classDef aopnode fill:#fafafa,stroke:#555,stroke-dasharray:3 3,color:#333;
  class intake,ingestion,reasoning,output boundary;
  class S3IN,S3OUT,AUDIT store;
  class INC2,CRED,ELIG,ESC aopnode;
```

*Fig. W1.1 — Mortgage Insurance Underwriting Agent full architecture. Every stage executes within your network boundary. The AOP governs all three reasoning domains simultaneously. Confidence scoring is the last step before output assembly — it cannot be bypassed.*

***

## 3 — Document intelligence layer

The MI workflow processes document types and volumes that are not handled by general-purpose document processing tools. A single loan file is routinely a few hundred pages across employer-specific pay stubs, bank-statement tables, tri-merge credit, tax-return schedules, and mixed-quality scans. The document intelligence layer is built specifically for that ecosystem.

***

### 3.1 Multi-pass OCR pipeline

MI loan files routinely include:

* Bank statements with institution-specific table layouts, running balances, and scanned or PDF-print artefacts
* Employer-specific pay stubs — checkboxes, earning codes, YTD columns, and handwritten annotations on the same page
* W-2 and 1040 packages with multi-schedule table grids
* Tri-merge credit reports with dense multi-column tradeline tables
* Handwritten or wet-ink letters of explanation and borrower certifications
* Appraisal PDFs that mix narrative, comparable grids, and photographs

The multi-pass OCR pipeline processes these document types through three sequential passes before any content is passed to the reasoning layer:

| pass                                | purpose                                                                                                                                                                                                                                                                           | output                                                                                                       |
| ----------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| **Pass 1 — Layout Analysis**        | Detect document structure: headers, transaction tables, form fields, handwritten regions, signature zones, and photo plates. Classify each region by content type before attempting extraction.                                                                                   | Region map with content-type labels and geometric coordinates                                                |
| **Pass 2 — Targeted Extraction**    | Apply model-specific extraction to each region type. Printed text, handwriting, bank-statement tables, pay-stub grids, and credit tradeline columns each receive a specialized extraction model.                                                                                  | Per-field extracted values with per-field OCR confidence scores                                              |
| **Pass 3 — Confidence Arbitration** | Where multiple extraction models produced competing outputs for the same field — particularly YTD earnings appearing on both a pay stub and a W-2, or a liability appearing on credit and the 1003 — the arbitration layer resolves the conflict or marks the field as uncertain. | Final extracted values with arbitrated confidence scores; uncertain fields marked with source ambiguity flag |

<Warning>
  A field marked as uncertain by the OCR pipeline is never silently passed to
  the reasoning layer as if it were a high-confidence extraction. The agent
  explicitly acknowledges the extraction uncertainty in its evidence citations
  and applies the appropriate confidence suppression to any finding that depends
  on that field. If the uncertain field is a material underwriting dimension —
  e.g., YTD income, a large deposit, representative FICO, or appraised value —
  the suppression is sufficient to trigger a requirements flag for the specific
  document and field.
</Warning>

***

### 3.2 Multi-document context assembly

The MI loan file is not a single document — it is typically fifteen to forty separate files totalling one hundred to several hundred pages. General-purpose LLM tools cannot process this volume in a single inference context, and processing documents sequentially loses the cross-document relationships that are the most valuable signal in MI underwriting.

The context assembly layer constructs a **unified case context** that makes all documents simultaneously available to the reasoning layer:

<CardGroup cols={2}>
  <Card title="Intelligent Context Prioritisation" icon="filter">
    Rather than concatenating all documents in full, the assembly layer identifies the highest-information-density content from each document and prioritises it into the reasoning context. Full document text is available for citation; the reasoning layer operates on the structured, prioritized representation.
  </Card>

  <Card title="Entity Resolution Across Documents" icon="link">
    Named entities — borrowers, employers, the subject property, creditors, and
    the appraiser — are resolved to unified identities across all documents before
    reasoning begins. This means the agent can immediately identify that "J.
    Rivera" on the pay stub and "Jordan Rivera" on the URLA refer to the same
    person, rather than treating them as separate entities.
  </Card>

  <Card title="Cross-Document Fact Graph" icon="diagram-project">
    Material facts extracted from multiple documents — income, liabilities,
    occupancy, property value, FICO, LTV — are assembled into a structured fact
    graph that explicitly tracks which documents support, contradict, or are
    silent on each fact. The reasoning layer operates on this graph, not on raw
    document text.
  </Card>

  <Card title="Third-Party Feed Integration" icon="plug">
    Structured feeds from credit bureaus, AUS (DU / LPA), employment verification, flood determination, and fraud / early-warning services are normalized into the same unified schema as the submitted documents. The agent treats third-party data and borrower-submitted documents as two independent evidence streams and explicitly identifies any divergence between them.
  </Card>
</CardGroup>

***

### 3.3 Third-party data normalisation

Third-party data sources for MI cases arrive in a variety of formats — credit returns proprietary multi-bureau XML, DU / LPA return structured findings requiring interpretation against the AUS finding glossary, employment verification returns insurer-specific schemas, and flood certificates vary by provider.

The normalisation layer converts all of these feeds into the unified MI case schema before any reasoning step:

```json theme={null}
{
  "third_party_data": {
    "credit": {
      "bureaus": ["equifax", "experian", "transunion"],
      "representative_score": 742,
      "score_method": "middle_of_three",
      "tradeline_count": 18,
      "inquiries_last_90_days": 2,
      "public_records": [],
      "extraction_confidence": 0.97
    },
    "aus": {
      "engine": "DU",
      "recommendation": "Approve/Eligible",
      "findings": [
        {
          "code": "GDOC",
          "category": "documentation",
          "significance": "condition"
        },
        {
          "code": "DTI",
          "category": "capacity",
          "significance": "informational"
        }
      ],
      "cross_reference_status": "aligned",
      "discrepancy_detail": null
    },
    "employment_verification": {
      "employer": "Northlake Regional Medical Center",
      "hire_date": "2019-03-11",
      "ytd_earnings": 68420,
      "status": "active",
      "disclosed_on_urla": true
    }
  }
}
```

<Note>
  The cross-reference status field in the normalized third-party schema is
  computed by the ingestion layer, not the reasoning layer. This means the agent
  enters the reasoning phase already aware of structural discrepancies between
  AUS findings, credit, VOE, and submitted documents — it does not discover them
  mid-reasoning. This architecture prevents the reasoning layer from anchoring
  on the loan file before evaluating third-party data.
</Note>

***

## 4 — Confidence engine calibration for MI

The confidence engine architecture is described in [the Confidence Engine reference](/agents/confidence/engine). This section covers the MI-specific elements: how the five signal domains are weighted for this workflow, and how the engine is validated against an MI-specific synthetic dataset before any model version is certified for production.

***

### 4.1 MI signal domain weights

The five confidence signal domains carry different weights in the MI composite than they would in a different workflow agent. The weight set is calibrated to reflect the relative evidential importance of each domain in mortgage insurance underwriting specifically.

| signal domain                   | MI weight | rationale                                                                                                                                                                                                                                                                            |
| ------------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Cross-Evidence Consistency**  | 30%       | Income, liabilities, occupancy, and property value are stated on the URLA and independently evidenced by pay stubs, W-2s, bank statements, credit, and the appraisal. Discrepancies between stated and evidenced facts are the highest-value underwriting signal in MI.              |
| **Document Extraction Quality** | 25%       | MI packets include some of the most challenging document types in insurance: bank-statement tables, employer-specific pay stubs, tri-merge credit, and handwritten letters of explanation. Extraction quality directly determines whether the material facts are recoverable at all. |
| **AOP Coverage**                | 20%       | MI cases regularly surface overlay, occupancy, and income-type combinations that lie at the edge of AOP coverage. AOP gaps in MI are more likely to be genuinely ambiguous than in simpler workflows and carry a higher escalation burden.                                           |
| **Inference Quality**           | 15%       | The MI reasoning chain is multi-step — income annualisation, DTI construction, LTV math, overlay screening, delegated-authority routing. Model uncertainty accumulates across those passes. This domain captures that accumulated uncertainty.                                       |
| **External Data Alignment**     | 10%       | AUS, credit, and VOE alignment is a critical signal but is treated as a detractor amplifier rather than a standalone domain — an AUS discrepancy alone does not resolve the case, but significantly amplifies the weight of any corroborating inconsistency in the submitted file.   |

***

### 4.2 MI-specific detractors

In addition to the standard detractor taxonomy, the MI workflow includes several detractors that are specific to this product line and do not appear in other agent configurations.

<CardGroup cols={2}>
  <Card title="Unsourced Large Deposit" icon="building-columns">
    A deposit on a bank statement exceeds the AOP large-deposit threshold and cannot be sourced to payroll, a documented asset transfer, or a gift letter in the file. This detractor fires regardless of whether the remaining assets still clear the reserve requirement — unsourced funds are an independent eligibility finding.
  </Card>

  <Card title="Income Document Reconciliation Break" icon="arrows-left-right">
    Annualised YTD pay-stub earnings, W-2 Box 1, VOE YTD, and URLA stated income
    cannot be reconciled within the AOP tolerance. This detractor carries a high
    severity weight because MI pricing and eligibility are directly tied to
    qualifying income.
  </Card>

  <Card title="Omitted Liability" icon="credit-card">
    A tradeline on the tri-merge with a monthly payment above the AOP materiality
    floor does not appear on the URLA or 1008. The agent does not assume the
    borrower will pay it off at closing unless a documented payoff is in the file.
  </Card>

  <Card title="Delegated Authority Breach" icon="scale-balanced">
    The case characteristics — FICO, LTV, DTI, occupancy, property type, exception
    count — place it at or beyond the delegated authority matrix cell. Cases at
    the boundary receive a confidence suppression and are flagged Refer
    Non-Delegated before disposition.
  </Card>

  <Card title="AUS–File Divergence" icon="robot">
    DU / LPA recommendation is Approve/Eligible but the extracted file does not
    support the AUS findings (or the reverse). This detractor is specific to the
    intersection of AUS and the document file and does not fire from either source
    independently.
  </Card>

  <Card title="Occupancy Inconsistency" icon="house">
    Declared occupancy is inconsistent with credit address history, URLA residence pattern, or appraisal occupancy indicators. Occupancy misrepresentation is treated as an eligibility finding, not a documentation nit.
  </Card>
</CardGroup>

***

## 5 — MI benchmark results

The agent is evaluated against an MI-specific synthetic validation dataset before any model version is certified for production. This section describes the validation methodology, the dataset, and the certified benchmark results for the current production model.

***

### 5.1 Validation methodology

The MI synthetic validation dataset is constructed by Layerup's underwriting implementation team to represent the full distribution of case complexity encountered in production mortgage insurance underwriting. Cases in the dataset are assigned ground-truth labels by experienced MI underwriters — not derived from model outputs.

The dataset is structured to deliberately stress every agent capability:

| dataset segment                      | construction method                                                                                               | purpose                                                                                             |
| ------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| **Clean cases**                      | Synthetically generated complete loan files with no deliberate defects                                            | Establish baseline precision; agent should resolve these at high confidence                         |
| **OCR-degraded cases**               | Clean cases with intentionally degraded document quality (rotation, noise, reduced DPI, handwriting substitution) | Validate multi-pass OCR pipeline under realistic document quality conditions                        |
| **Cross-document discrepancy cases** | Cases with deliberate income, liability, or occupancy inconsistencies inserted across document pairs              | Validate cross-evidence consistency detection; agent must surface all inserted discrepancies        |
| **AOP gap cases**                    | Cases with income type, occupancy, or overlay combinations not covered by the AOP                                 | Validate that the agent surfaces gaps rather than extrapolating; no recommendation should be issued |
| **AUS / credit divergence cases**    | Cases where AUS or credit data reveals facts not supported by the submitted file                                  | Validate third-party data normalisation and discrepancy detection                                   |
| **Delegated-boundary cases**         | Cases synthetically placed at delegated authority matrix boundaries                                               | Validate boundary detection and appropriate confidence suppression                                  |

Before any AOP version or model version is promoted to production, the agent must pass all six benchmark checks listed in Section 5.2. Failure on any single check blocks promotion.

***

### 5.2 Certified benchmark results — current production model

The following benchmarks are the certified results for the current production model version against the MI synthetic validation dataset. Results are expressed as the mean and interquartile range (IQR) across three independent validation runs.

**Recommendation accuracy** — overall alignment between agent recommendation and ground-truth underwriter disposition:

| decision type                     | precision | recall    | F1        |
| --------------------------------- | --------- | --------- | --------- |
| Approve as Applied                | 99.0%     | 99.2%     | 99.1%     |
| Approve with Conditions / Overlay | 98.5%     | 98.8%     | 98.6%     |
| Refer Non-Delegated               | 99.3%     | 99.1%     | 99.2%     |
| Defer Pending Requirements        | 98.8%     | 99.0%     | 98.9%     |
| Escalate                          | 99.5%     | 99.4%     | 99.4%     |
| Decline / Ineligible              | 99.1%     | 99.0%     | 99.0%     |
| **Overall weighted**              | **99.0%** | **99.1%** | **99.0%** |

**Cross-document discrepancy detection** — on the discrepancy segment of the validation set:

| discrepancy type                               | detection rate | false positive rate |
| ---------------------------------------------- | -------------- | ------------------- |
| Income discrepancy (W-2 vs. pay stub vs. URLA) | 99.4%          | 0.5%                |
| Unsourced large deposit                        | 99.1%          | 0.4%                |
| Omitted liability (credit vs. 1003)            | 99.2%          | 0.4%                |
| Occupancy inconsistency                        | 98.6%          | 0.7%                |
| AUS–file divergence                            | 98.9%          | 0.5%                |
| Delegated authority boundary                   | 99.3%          | 0.3%                |

**Confidence score calibration** — Brier score and Expected Calibration Error (ECE) measure how accurately the composite confidence score predicts actual recommendation accuracy:

| metric                                 | MI result | interpretation                                                                                                                                                    |
| -------------------------------------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Brier Score                            | 0.009     | Lower is better; 0 = perfect calibration. A score of 0.009 indicates the confidence score is a highly reliable predictor of recommendation accuracy.              |
| Expected Calibration Error (ECE)       | 0.8%      | The average gap between the stated confidence and the empirical accuracy across the confidence range.                                                             |
| Cases correctly escalated at threshold | 99.5%     | Proportion of validation cases that should have received human review (ground truth) that did receive hard escalation from the confidence engine.                 |
| Cases incorrectly auto-resolved        | 0.1%      | Proportion of validation cases where the agent produced a recommendation at or above the high threshold but the ground-truth disposition was Escalate or Decline. |

**OCR extraction accuracy** — on the OCR-degraded segment of the validation set:

| document type                       | field-level extraction accuracy | extraction floor hit rate |
| ----------------------------------- | ------------------------------- | ------------------------- |
| Bank statements (variable layout)   | 98.8%                           | 0.8%                      |
| Pay stubs (employer-specific)       | 98.9%                           | 0.7%                      |
| W-2 forms                           | 99.2%                           | 0.4%                      |
| Credit reports (tri-merge)          | 98.6%                           | 1.0%                      |
| Tax returns (1040 + schedules)      | 99.1%                           | 0.5%                      |
| URLA / 1003 (mixed print + wet ink) | 99.0%                           | 0.6%                      |
| Appraisal reports (tables + photos) | 98.7%                           | 0.9%                      |
| Handwritten letters of explanation  | 98.1%                           | 1.4%                      |
| Standard printed PDFs               | 99.8%                           | 0.1%                      |

<Note>
  "Extraction floor hit rate" is the proportion of pages where OCR confidence
  falls below the configured extraction floor threshold, triggering an
  uncertainty flag rather than passing the extracted value to the reasoning
  layer. A higher floor hit rate on handwritten letters of explanation and
  tri-merge credit reflects a conservative extraction policy — the agent prefers
  to flag and require human verification over silently passing low-confidence
  income, liability, or credit-event extractions.
</Note>

**AOP gap handling** — on the AOP gap segment of the validation set:

| gap scenario                                               | correct gap surface rate | hallucination rate |
| ---------------------------------------------------------- | ------------------------ | ------------------ |
| Income type outside AOP documentation hierarchy            | 100%                     | 0%                 |
| Occupancy / property-type combination with no overlay rule | 99.7%                    | 0%                 |
| Multi-state regulatory conflict                            | 99.2%                    | 0%                 |
| Intra-AOP rule conflict                                    | 99.4%                    | 0%                 |

<Warning>
  A hallucination rate of 0% in the AOP gap segment means the agent produced
  zero instances where it issued a recommendation by extrapolating beyond its
  AOP coverage. This is a hard architectural invariant: the agent is not
  permitted to issue a recommendation for a case dimension where no AOP rule
  applies. The zero result is expected and is enforced by the output assembly
  layer — it validates AOP coverage before assembling any recommendation output.
</Warning>

***

### 5.3 Score distribution under production conditions

The certified baseline confidence score distribution across the validation set establishes the reference against which all production AOP updates are evaluated. Any AOP promotion that shifts the distribution beyond the configured tolerance band is blocked by the CI/CD test harness.

| confidence tier                               | validation set distribution | production target range |
| --------------------------------------------- | --------------------------- | ----------------------- |
| Auto-Resolve (≥ high threshold)               | 56.8%                       | 52–64%                  |
| Soft Review (mid threshold to high threshold) | 30.1%                       | 26–36%                  |
| Hard Escalation (\< low threshold)            | 13.1%                       | 10–18%                  |

<Note>
  The production target range is intentionally wide enough to accommodate
  genuine variation across customer AOP configurations — delegated vs.
  non-delegated mix, overlay tightness, income-type mix — while being narrow
  enough to detect implausible distributions. A distribution concentrated above
  70% auto-resolve would suggest a permissive AOP configuration that is not
  applying appropriate sensitivity to MI-specific risk signals. A distribution
  with over 25% hard escalations would suggest an AOP that is not covering the
  case population adequately.
</Note>

***

## 6 — Structured output schema

Every MI case processed by the agent produces a single structured JSON output payload. The following describes the top-level schema and the MI-specific fields.

```json theme={null}
{
  "case_id": "MI-2024-110482",
  "certificate_number": "MI-001-482910",
  "processed_at": "2024-12-04T14:37:22Z",
  "processing_time_seconds": 912,

  "borrower": {
    "name": "Jordan Rivera",
    "credit_score_representative": 742,
    "occupancy": "primary",
    "income_type": "W-2 employment"
  },

  "loan": {
    "purpose": "purchase",
    "loan_amount": 468000,
    "appraised_value": 520000,
    "purchase_price": 515000,
    "ltv": 90.87,
    "cltv": 90.87,
    "dti": 41.2
  },

  "income_analysis": {
    "income_type": "W-2 employment",
    "qualifying_monthly_income": 8750,
    "prior_year_income": 102400,
    "ytd_annualised": 105180,
    "income_trend": "stable",
    "income_sufficiency_score": 0.96,
    "reconciliation_breaks": [],
    "financial_requirements": []
  },

  "credit_analysis": {
    "representative_score": 742,
    "score_method": "middle_of_three",
    "omitted_liabilities": [],
    "derogatory_events": [],
    "inquiries_last_90_days": 2,
    "letter_of_explanation_adequacy": "not_required"
  },

  "eligibility_analysis": {
    "guideline_result": "eligible",
    "overlay_result": "eligible",
    "delegated_authority": "within_authority",
    "aus_alignment": "aligned",
    "appraisal_sufficiency": "adequate",
    "conditions": []
  },

  "ai_recommendation": {
    "decision": "Approve as Applied",
    "confidence_score": 94,
    "conditions_recommended": [],
    "overlays_applied": [],
    "requirements": [],
    "escalation_required": false,
    "confidence_signal_breakdown": {
      "composite_score": 94,
      "signals": [
        {
          "domain": "cross_evidence_consistency",
          "domain_score": 96,
          "weight": 0.3
        },
        {
          "domain": "document_extraction_quality",
          "domain_score": 93,
          "weight": 0.25
        },
        { "domain": "aop_coverage", "domain_score": 100, "weight": 0.2 },
        { "domain": "inference_quality", "domain_score": 91, "weight": 0.15 },
        {
          "domain": "external_data_alignment",
          "domain_score": 88,
          "weight": 0.1
        }
      ]
    }
  },

  "evidence_citations": [
    {
      "finding": "YTD pay stub reconciles to W-2 within 2.7%",
      "sources": [
        {
          "document": "paystub_2024-11-15.pdf",
          "page": 1,
          "field": "ytd_gross"
        },
        {
          "document": "w2_2023.pdf",
          "page": 1,
          "field": "box_1_wages"
        }
      ]
    }
  ],

  "documents_processed": 22,
  "total_pages_reviewed": 186,
  "aop_version": "v1.3.0",
  "model_version": "layerup-mi-uw-v2"
}
```

***

## 7 — KPI impact model

The MI Underwriting Agent is designed to produce measurable, auditable improvements across the underwriting operations KPIs that matter most — including cost per loan file, cycle time, and referral discipline. The following describes the KPI model and the mechanisms by which each improvement is achieved.

***

### 7.1 Throughput & cycle time

<CardGroup cols={2}>
  <Card title="Case Throughput" icon="gauge">
    The agent runs in parallel across all queued loan files simultaneously — there is no single-threaded bottleneck. Throughput scales with compute configuration, not with headcount. Flow volumes that historically required hiring or overtime can be absorbed by scaling the agent's task configuration.
  </Card>

  <Card title="Cycle Time Compression" icon="clock">
    The agent completes the initial file review — including document extraction,
    income reconciliation, credit/DTI construction, guideline screening, and
    requirement generation — in minutes rather than hours. The human underwriter
    receives a fully prepared loan-file record rather than a raw document stack.
  </Card>

  <Card title="Requirement Latency Reduction" icon="list-check">
    The requirements list is generated at the end of the agent's first pass —
    before any human underwriter touches the file. Outstanding conditions reach
    the lender or correspondent earlier, reducing the time files spend in pending
    status due to incomplete documentation.
  </Card>

  <Card title="In-Good-Order Rate Improvement" icon="circle-check">
    By generating targeted, document-specific requirements rather than generic follow-up requests, the agent reduces the round-trip count needed to resolve conditions. Correspondents receive precise instructions — "provide November YTD pay stub for employer Northlake Regional covering the period after the 2023 W-2" rather than "provide updated income docs."
  </Card>
</CardGroup>

***

### 7.2 Underwriting quality & consistency

<CardGroup cols={2}>
  <Card title="Consistent AOP Application" icon="equals">
    The agent applies the AOP with zero variance across every file. No file is processed differently because of the time of day, the underwriter's experience level, or workload pressure. Income inclusion rules, FICO / LTV / DTI thresholds, and delegated-authority boundaries are applied identically on every case.
  </Card>

  <Card title="Discrepancy Detection Rate" icon="magnifying-glass">
    The cross-reference between borrower-submitted documents and third-party data
    (credit, AUS, VOE) surfaces omitted liabilities, unsourced deposits, and
    AUS–file breaks that human review of submitted documents alone cannot reliably
    identify. The agent's validated income-reconciliation detection rate exceeds
    99% on the benchmark dataset.
  </Card>

  <Card title="Intra-AOP Conflict Surface" icon="arrows-split-up-and-left">
    As the agent processes files against the AOP, it surfaces situations where
    guideline and overlay clauses conflict — inconsistencies that may have existed
    in the underlying SOP for years without being identified because individual
    underwriters resolved them informally.
  </Card>

  <Card title="Decision Auditability" icon="receipt">
    Every recommendation the agent issues includes a full evidence citation chain from raw document to extracted fact to reasoning step to decision. Audit and compliance review of any historical file can trace the complete reasoning without re-running the agent or relying on underwriter notes.
  </Card>
</CardGroup>

***

### 7.3 Human underwriter productivity

The agent is designed to make the underwriter's remaining work higher-value, not to create additional review burdens. Files that reach a human underwriter from the agent queue arrive with:

* A structured recommendation with income verification, DTI construction, and guideline / overlay screening already completed
* A prioritized requirements list with document-specific, targeted conditions
* A confidence signal breakdown that tells the underwriter exactly which evidence domain is weakest and why
* Full evidence citations for every finding, so the underwriter can immediately navigate to the relevant document and page

The net effect is that the underwriter's cognitive load on a file prepared by the agent is concentrated on the genuinely ambiguous elements — the ones the confidence engine surfaced as requiring expert judgment — rather than spread across the full document stack.

| underwriter activity             | before agent         | with agent                                                                         |
| -------------------------------- | -------------------- | ---------------------------------------------------------------------------------- |
| Document review (full stack)     | 40–80 min per file   | 0 — agent produces indexed, cited findings                                         |
| Income documentation review      | 20–40 min per file   | 0 for reconciled files; focused review of flagged discrepancies only               |
| Credit / DTI construction        | 15–25 min per file   | 0 — agent normalizes tri-merge and constructs DTI before the file reaches UW       |
| Guideline / overlay screening    | 10–20 min per file   | 0 for in-matrix files; 8 min review for overlay or delegated-boundary flags        |
| Condition / requirement drafting | 10–20 min per file   | 0 — agent generates targeted requirement list for underwriter approval             |
| **Focused case judgment**        | Residual after above | **Primary underwriter activity — exercised on the files the agent cannot resolve** |

***

### 7.4 Risk & compliance KPIs

<CardGroup cols={2}>
  <Card title="False Approval Rate" icon="shield-x">
    The confidence engine's calibrated escalation logic ensures that files the agent is not entitled to approve do not auto-resolve. The validated false approval rate in the MI benchmark set — cases where the agent issued an Approve recommendation but the ground-truth disposition was Escalate or Decline — is 0.4%.
  </Card>

  <Card title="Escalation Capture Rate" icon="arrow-up-right">
    99.4% of validation cases with a ground-truth escalation or
    non-delegated-referral disposition received the corresponding hard route from
    the agent. Escalations are never suppressed by the output assembly layer — the
    escalation flag is written before the recommendation payload is assembled.
  </Card>

  <Card title="Regulatory Exam Readiness" icon="file-check">
    The structured, evidence-cited output payload and the immutable CloudWatch
    audit trail give your compliance team a complete, immediately accessible
    record for any file under regulatory examination — without requiring
    underwriter recall or document reconstruction.
  </Card>

  <Card title="SOP Enforcement Consistency" icon="book-open-check">
    The agent applies your AOP-encoded SOP identically across 100% of files. Enforcement consistency across flow vs. bulk, delegated vs. non-delegated, and state overlays is a direct output of AOP-governed operation — not a function of underwriter training or manager oversight.
  </Card>
</CardGroup>

***

## 8 — Rollout & human-in-the-loop model

The MI Underwriting Agent is deployed through a structured rollout model that preserves human decision authority at every stage. The agent's role expands progressively as confidence in its calibration is established through real production data.

```mermaid theme={null}
flowchart LR
  P1["Phase 1 — Shadow Mode<br/>Agent runs on every file<br/>Underwriter makes all decisions<br/>Agent output visible but advisory only"]
  P2["Phase 2 — Assist Mode<br/>Agent output is primary review surface<br/>Underwriter reviews all recommendations<br/>Accepts, modifies, or overrides each"]
  P3["Phase 3 — Auto-Resolve (High Confidence)<br/>Files above high confidence threshold<br/>auto-resolve without mandatory review<br/>Underwriter reviews queue of flagged files"]
  P4["Phase 4 — Full Deployment<br/>Auto-resolve + soft review queue<br/>Hard escalations route to senior underwriter<br/>Ongoing benchmark monitoring active"]

  P1 -->|"Calibration validated<br/>against shadow outcomes"| P2
  P2 -->|"Recommendation alignment<br/>confirmed > threshold"| P3
  P3 -->|"Auto-resolve error rate<br/>confirmed < tolerance"| P4

  classDef phase fill:#fafafa,stroke:#111,color:#111;
  class P1,P2,P3,P4 phase;
```

*Fig. W1.2 — MI rollout phases. Each phase transition is gated by a measurable calibration check. No phase transition is automatic — your underwriting governance team approves each gate.*

<Note>
  Human decision authority is preserved unconditionally throughout Phases 1 and
  2, and for all hard-escalation and delegated-authority-breach cases in Phases
  3 and 4. The agent does not make a binding underwriting decision at any phase
  — it produces a recommendation. Your underwriting governance team configures
  the confidence threshold at which recommendations are treated as actionable
  without mandatory individual review.
</Note>
