Parsewise achieves SOTA on Databricks OfficeQA benchmark – results

The Risk of SamplingHow measuring the narrative record reduces ambiguity in complex transactions

Maximilian Hofer3 September 2026

Field of pale dots with small blue and pink samples highlighted

TLDR: The structured record shows only part of the risk in a complex transaction. The rest sits in narrative records that experts can sample but cannot review in full. Measuring the narrative record can reduce ambiguity and lead to sharper pricing. Finally, we describe practical considerations to build measures that are precise, comparable, stable, and traceable.

Examples for transaction data
Structured recordNarrative record
Private equity transactionse.g. leveraged buyouts, carve-outs, co-investments
  • Audited financials
  • Customer cube
  • Management accounts
  • Customer references
  • Management presentations
  • Contracts
Insurance transactionse.g. loss portfolio transfers, sidecars
  • Loss runs
  • Reserve triangles
  • Reinsurance recoveries
  • Adjuster notes
  • Medical records
  • Legal correspondence

In any transaction, one side discloses information and the other assesses it. Part of that information sits in the structured record: standardised and complete, giving you a distribution of outcomes if you were to execute the transaction. The rest sits in the narrative record, which is neither standardised nor complete, typically used by humans sampling data points.

Illustrative figure

Two Records, One Transaction

The Structured RecordRead in full100%Standardised and completeThe Narrative Recordread by a specialistnot openedRead by sample~3%Neither standardised nor complete

Proportions are illustrative. The structured record is small enough to read completely. The narrative record is not.

Measurement here means reading, transforming, and analysing in support of decision-making. The limitation lies in the human capacity to fully measure the narrative record in a limited timeframe.

How the bid price for an LPT gets built

  • Start with the structured record: the loss history gives you a distribution of possible portfolio losses across all claims. Compute the expected loss range.
  • Add your margin for profit and risk, which is the business case of taking that distribution onto your books, plus a buffer.
  • Add an ambiguity premium based on the narrative record. You sample data points to make the initial distribution you got from the loss history more accurate. The pessimistic end of the range is the price you bid.

The idea of the ambiguity premium builds on Epstein and Schneider (2008). Think of it as compensating you for not knowing which distribution you are holding. The risk margin, on the other hand, compensates you for holding a distribution you can define. Value at Risk (VaR) and Tail Value at Risk (TVaR), for example, quantify the risk of that distribution.

The ambiguity premium sits a layer above those measures. A useful intuition is entropy: the sampled narrative record leaves many distributions consistent with what you have read, and every file you measure rules some of them out. Sampling the narrative record shrinks the ambiguity premium. Fully measuring it shrinks it further.

Below are visualizations for both cases, where sampling over- or understates the true risk. In practice, sampling typically understates the true risk because unopened small claims today sit in the thin tail.

Illustrative figure

How the bid price gets built

The structured record sets the starting distribution. The narrative record revises it. How far it revises it is what a sample cannot tell you.

05 / 05
Completed bid price built from the structured and measured narrative records400450500550600650700Ultimate portfolio loss, $mStructured record onlyPlus the measured recordExpected loss$500mPlus margin$545mBid after sampling$595mProfit and risk margin$45mAmbiguity premium$50mThe bid rises by $33mBid after measurement$628mBuild-up of a bid price from the structured and narrative records
05

Measure the narrative record in full and the range collapses to one revised distribution. Here it sits above what the sample implied, so you bid higher or you walk away. You give up a portfolio you would have written at the wrong price.

All figures are invented for illustration and describe no actual portfolio. The amber curves show the range of loss distributions that fit a sample of the narrative record. They are not the range of outcomes. Measurement does not change what the claims will cost.

If the true distribution sat below what sampling suggested, the narrative record held less risk than the sample implied. You can bid lower and still hold your margin, which makes you more likely to win the portfolio. If the true distribution sat above what sampling suggested, the narrative record held more risk than the sample implied. You bid higher, or you walk away, and you avoid slipping on hidden banana skins that would have surfaced years later as adverse development. In both cases, you priced what was actually there rather than what a few files implied.

Let's run a back-of-the-envelope example. Consider a book with an expected loss of 500 million and an ambiguity premium of 8-10%. 40-50 million of your price depends on the narrative record. If you compress half of it, you have 20 to 25 million of pricing headroom on a single transaction, to spend on winning the deal or to protect your margin. The figures are illustrative, but the order of magnitude is realistic.

Let's look at an LPT example

Take an LPT of ten thousand bodily injury claims, some closed, some open. The structured record is in good shape. You have paid and incurred by claim, reserve movements, dates of loss, jurisdictions, and enough development history to build a triangle. That gives you a distribution of portfolio losses and an expected loss range.

The narrative record for the same portfolio includes a few hundred thousand pages of adjuster notes, medical reports, and correspondence with defence counsel. Records have been written to manage claims, not for a potential buyer to review them. Therefore, specialist review timelines often exceed deal timelines. In a competitive process, you might only have two weeks and have to take shortcuts to meet non-binding offer (NBO) deadlines. Measurement that runs in hours changes what you can price confidently and how quickly.

A claims specialist reads a file to get risk cues such as step-laddering reserves, injury treatment gaps, and static claims. This is a measurement process. It takes the specialist an hour or more per claim, so she samples what look like expensive claims.

The challenge with that is that the claims becoming expensive in the future are usually the ones that change trajectory. For example, a soft tissue claim that becomes a surgical claim or a return to work that does not persist. When that shift begins, the claim is still small, and the structured record looks benign, so it is not included in the sample. However, the narrative record already shows the pattern.

Illustrative figure

Sampling follows outdated signals

05 / 05

A reserve-weighted review focuses on existing developments, rather than future developments.

Ten thousand bodily injury claims, by reserve todayclaim in the bookread in the samplenotes already show a change of trajectoryReserve-weighted sample~300 claims1k10k100k1mIncurred reserve today, log scaleStill small today, so the sample passes over themOne of those claims, followed throughClaim notes contain the signal for eventual reserve changes.Claim notesTreatment escalatesAttorney instructedReturn to work failsIncurred reserveReserve strengthened012243648months24 months between the first note and the reserve move
05

The reserve moves much later. Two years after the first note, the incurred is strengthened. The sample had already closed.

Illustrative. A claim is still small at the moment its trajectory changes, so a reserve-weighted sample passes over it. The notes already carry the signal.

The ambiguity premium is covering LPT buyers for exactly this case: the non-zero probability that some claims sit quietly in the portfolio and look calm on the surface, but contain material signal in the narrative record.

How to use the narrative record for better pricing

The intuitive response is to instruct an LLM to read everything in the narrative record. However, reading everything is not the same as measuring anything, and an LLM set to maximum reasoning effort will produce a lot of text, but very little output that can be priced.

I ran into this problem shape in the financial services context. One of my PhD papers studied text-based risk disclosure across 2,532 IPOs, covering firms whose value was largely in intangible assets that investors could not readily measure. The setting is useful because IPO underpricing, the first-day return when a new company comes to market, is a return that the market pays investors for holding risk.

The result was that aggregate disclosed risk explained nothing. More text, more risk factors, and more pages showed no association with first-day return at all. At the same time, one precisely constructed variable explained roughly ten percentage points of first-day return. That variable was technology risk, defined narrowly enough to mean a single thing. That association was far weaker for firms already holding granted patents, a mechanism that makes an otherwise ambiguous asset legible to the public. In short, the narrative record of IPO risk disclosures contains signal that matters for pricing.

The same pattern shows up in the insurance sector. The Casualty Actuarial Society published work in 2026 on converting narrative claims information into structured actuarial variables, which is the same construction applied in the insurance sector. The legacy market has also been solving the problem structurally. Aon's Lloyd's Legacy Report finds that ~56% of reserves transacted at Lloyd's since 2015 came from repeat sellers. On a second transaction with the same cedant, you have already read their files and know their reserving conventions, so the range of outcome estimates is narrower from the get-go. Execution certainty and an established working relationship explain part of that pattern, though likely not all of it.

The lesson for practitioners can be summarized in a simple sentence: A well-specified construct is priced, but an aggregate is not. The following four principles are useful to keep in mind when moving from sampling to complete measurement of narrative records:

  • Decompose before you aggregate. A single risk score that works across lines of business and geographies is likely explaining little true risk. Instead, measure well-specified constructs separately (e.g., treatment escalation, return-to-work trajectory). Claim specialists recognise these constructs, which builds trust and unlocks adoption.
  • Count discrete events. A file-level score cannot distinguish one mention of escalation from eleven spread across eighteen months. Documents can vary in structure and length, but the risk measure is often independent of those technical factors.
  • Normalise against comparables. A “risk score of 7.2” is just a number, but “1.8 standard deviations above jurisdiction, line of business and accident year peers” is a variable you can directly use for pricing. Documentation practices drift over time and across jurisdictions and TPAs, which makes an unnormalised measure confuse that drift as risk.
  • Ensure stability and traceability. Log LLM versions, keep a frozen evaluation set, and assess any new model update against it. Keep every value traceable to the sentence that produced it, so an investment committee or an auditor can follow the chain from a number back to the narrative record it came from.
Completed scorecard05 / 05

Reference card

Measuring the narrative record

01

Decompose before you aggregate

Measure named constructs rather than a single score.

02

Count discrete events

Eleven mentions vs a single one are different.

03

Normalise against comparables

Compare against peers, not a raw score.

04

Keep it stable and traceable

Trace every value back to its sentence.

REFERENCE CARDMeasuring the narrative record01Decompose before you aggregateMeasure named constructs rather than a single score.02Count discrete eventsEleven mentions vs a single one are different.03+1.8Normalise against comparablesCompare against peers, not a raw score.041.8Keep it stable and traceableTrace every value back to its sentence.From "The Risk of Sampling"parsewise.ai

Keep the completed scorecard as a reference.

Final thoughts

The constructs you specify for pricing continue being useful post-bind. The same measures, run on the same portfolio each quarter, tell you which claims to prioritize while there is still time to act. Pricing and value creation are the same measurement problem observed at two different moments.

Reducing the risk of sampling with complete measurement of the narrative record does not remove the transaction risk. The claims will cost what they cost. What changes is that one can make more of the eventual costs visible before binding rather than after.

The same dynamic appears in any transaction where one side has complete information, and the other approximates it from what is disclosed. What has changed is that the narrative record can now be measured in full rather than sampled, so the information gap between the two sides closes before the price is set rather than after. Closing that gap protects your margin on the transactions you win and reduces the volatility of the ones you have completed.