Industry Insights

AI Insurance Fraud Detection That Keeps Claims Data In-House

Malavika Kumar
Director of Product Marketing
Published September 14, 2026

The demo always works. A vendor loads a sample of your closed claims into their platform, runs the model, and shows you 3 fraud rings your SIU missed. Then someone from security asks where the claims data went to produce that result and the meeting gets quiet. That question is the whole problem with AI insurance fraud detection as it's usually sold. 

‍

The models are good. Some of them are very good. But most of them need to see your entire claims history to be good. And considering the potential costs associated with insurance fraud, your organization can’t afford a mediocre solution. Deloitte estimates that roughly 1 in 10 property and casualty (P&C) claims is fraudulent, costing carriers about $122 billion a year. With P&C accounting for 40% of the insurance industry's total fraud losses.

‍
The standard way to give them that is to copy it somewhere you don't control. For a carrier operating under state privacy law, model risk guidance, and a dozen reinsurance and TPA agreements, that copy is a liability before it's an asset.

‍

For that reason, we want you to know there's a better architecture. Bring the model to the claims, not the claims to the model. Here's what that looks like and how to get it running without stalling the claims queue.

‍

Why does fraud detection need the whole claims history?

Fraud isn't a property of a claim. It's a property of a pattern across claims. And the pattern only shows up when the model can see far enough. A single water damage claim looks fine. The same claimant filing a water damage claim on 3 properties in 4 years, each with the same restoration contractor, and each just under the inspection threshold, is a different story. 

‍

A staged accident looks like an accident until you notice the passenger was a passenger in 2 other accidents with 2 other drivers who share a chiropractor. None of that is visible inside one claim file. All of it is visible across the book.

That's why AI insurance fraud detection is hungry in a way that, say, a document summarizer isn't. It needs the claims system, the policy admin system, the vendor and provider directories, the payment history, and often the SIU case notes. It needs those records in a consolidated view that spans over years. 

‍

The trouble is that the same data that makes the model useful is the data you're least able to hand over. Claims files hold medical records, financial details, police reports, and personal identifiers for claimants who never agreed to have their information trained on by a third party. Copying that into a vendor environment is a legal and contractual decision before it's a technical one.

‍

Where do claims go when the fraud model lives in a vendor cloud?

Follow the data on a typical deployment and the path is longer than the sales deck suggests. Claims are extracted from your core system, usually nightly, and land in a staging area. They're transformed to the vendor's schema, which means someone has to map your claim codes, loss types, and party roles to theirs. 

‍

Then the transformed data moves to the vendor's cloud, where the model scores it and pushes alerts back. Somewhere there's a data processing agreement that says the vendor may use "aggregated and anonymized" data to improve its models. But nobody asks what anonymized means for a claims record with a VIN, a street address, and a treating physician in it.

‍

Every one of those hops is a place where a regulator, a plaintiff's attorney, or a reinsurer can ask a question you'd rather not answer. And state insurance departments have gotten specific about this. The NAIC's model bulletin on insurer use of AI, adopted by a growing list of states, expects carriers to document where data used by AI systems comes from, how it's governed, and who's accountable for the outputs. 

‍

New York's Department of Financial Services went further for underwriting and pricing, requiring insurers to show that external data and models don't produce unfair outcomes. Fraud detection isn't exempt from that logic just because the target is a bad actor. A false positive is a policyholder whose legitimate claim gets delayed, and that's a market conduct issue.

‍

Every copy of claims data you make is a copy you have to govern. And a fraud program that quietly creates a second copy of your entire book has doubled your exposure before it's caught a single ring. Our buyer's guide to AI data security walks through the questions to ask before that copy gets made.

‍

What does AI insurance fraud detection look like regarding data?

Instead of moving claims to the model, connect the model to the claims where they already sit. In practice, that means a governed layer inside your environment that knows how to read your claims system, your policy admin platform, your provider directory, and your SIU tools, in their native schemas but under your existing access controls. 

‍

The model, whether it's a graph-based ring detector, an anomaly model, or a language model reading adjuster notes, runs against that layer. Scores, alerts, and explanations flow into the claims workflow. Raw claims data doesn't leave.

‍

This is the pattern behind custom AI that doesn't require sharing your data. And it's what Unframe builds for insurers. Which is a layer that connects to policy admin, claims, CRM, and risk engines without migrating any of them. So anomaly detection can run across the full book in real time while the book stays home.

‍

Three things change when you do it this way:

‍

  1. The join problem gets solved once, in a form every future use case can reuse, instead of being rebuilt in a vendor's schema for each new tool.
  2. The security review gets shorter, because there's no new data store to assess. 
  3. You keep the option to run the whole thing on-premise or in a private cloud for the lines where that's required, without changing the architecture.

How do you catch fraud rings without a data lake?

The objection at this point is usually technical. Ring detection needs to compare millions of records across claims, providers, and payments, so surely all of those records have to be copied into one place first. They don't. The model needs to query across those systems as if they were one. It doesn't need a new database that holds a copy of everything. 

‍

Entity resolution, the work of deciding that "J. Alvarez" on a claim and "Jose Alvarez" on a provider invoice and "Alvarez Rehab LLC" on a payment are the same node in the graph, is a mapping task. It produces a relatively small index of entities and relationships. That index can live inside your perimeter, be refreshed as claims move, and be queried by the model without the underlying records ever being consolidated into a new database.

‍

The same applies to the unstructured half of the problem, which is where language models earn their place. Adjuster notes, medical narratives, police reports, and repair estimates carry most of the signal in a suspicious claim. And almost none of it is in a structured field. Document intelligence can read those in place, extract the entities and events that matter, and feed them into the graph. Our piece on AI for insurance data ingestion and multimodal analysis covers the mechanics.

‍

What you end up with is AI insurance fraud detection that can see the ring, explain why it's a ring, and hand the SIU a case file with the linked claims, providers, and documents already attached. What you don't end up with is a second copy of every claim you've ever paid.

‍

What do regulators expect from AI insurance fraud detection?

More than they used to, and the trend only runs one way. There are three expectations that come up in nearly every state framework:

‍

  • You have to be able to explain a flag. A score of 0.91 with no reasons attached is a problem when the claimant's attorney asks why the claim was delayed 6 weeks. 
  • You have to be able to show the data lineage. So documenting which records fed the decision, from which systems, and under which access rights. 
  • And you have to show human accountability. An SIU investigator has to be able to overrule the model, and the overrule has to be logged.

‍

An in-place architecture makes all 3 easier, not harder. Lineage is native, because the model reads from the systems of record rather than a transformed copy. Explanations can reference actual claim numbers and document pages. And human review sits inside the existing claims workflow instead of a separate vendor console nobody checks. 

‍

There's a commercial angle too. If a vendor is confident in its detection rates, it should be willing to price against them. A fraud program bought on outcome-based terms, where the fee tracks confirmed fraud identified rather than seats or volume, aligns the vendor with the SIU instead of with the number of alerts generated.

‍

Our guide to building AI guardrails without slowing down goes deeper on how to make the controls part of the system rather than a gate in front of it.


How do you roll out AI fraud detection without stalling the claims queue?

The fastest way to kill a fraud program is to route every flagged claim to an already overloaded SIU. Alert fatigue sets in by week 3 and the model gets blamed for a workflow problem. The trick is to start narrow and in shadow mode. 

‍

Pick a line of business and a fraud pattern. Ideally one where the SIU already has confirmed cases to validate against.

‍

Run the model in the background against live claims for 30 days without changing how anything is routed. Compare its flags to what investigators found on their own. That gives you a precision number you trust. And it tells you where to set the threshold before a single adjuster sees an alert.

‍

Then integrate the signal into the existing workflow rather than bolting on a new one. A flag should show up in the claim file the adjuster is already looking at, with the reasons and the linked claims attached. That way the decision to refer is one click and the decision to clear is logged. Our writeup on AI powered claims automation covers how the fraud signal fits into straight-through processing without becoming a bottleneck.

‍

Then you expand by pattern, not by tool. Once the first detector is producing confirmed cases, add the next pattern on the same connections and the same entity index. This is where the in-place architecture pays for itself. The second detector costs a fraction of the first because the expensive work of joining your systems was done once. If a vendor's plan for pattern two involves a second data extract and a second schema mapping, you've learned something about what you bought.

‍

Fraud is a pattern problem, and patterns need the whole book. The answer to that isn't to ship the book to whoever has the best model this quarter. It's to build one governed view inside your walls and let the models come to it. If you want to see that running against your own claims data, in your own environment, schedule a working session with us.

‍

FAQ

‍

Do we need to move our claims data to a vendor platform for AI fraud detection to work?

No. The models need a joined view of claims, policy, provider, and payment data, but that view can be built and queried inside your environment. Vendors that require a full extract are making an architectural choice, not a technical necessity.

‍

How does AI insurance fraud detection identify fraud rings?

By resolving entities across claims (claimants, providers, contractors, addresses, vehicles) into a graph and looking for relationship patterns that individual claims don't reveal, such as shared providers across unrelated claimants or repeated near-threshold claims. Language models add signal from adjuster notes and documents.

‍

Will AI fraud detection increase false positives and delay legitimate claims?

It can if thresholds are set on alert volume rather than validated precision. Running the model in shadow mode against confirmed SIU cases before it touches live routing, and integrating flags into the adjuster's existing view with reasons attached, keeps false positives manageable and auditable.

‍

What do regulators require for AI used in insurance fraud detection?

Expectations vary by state, but the common threads are explainable flags, documented data lineage, human oversight with logged overrides, and evidence that the system doesn't produce unfair outcomes for policyholders. In-place architectures make lineage and explanation easier to demonstrate.

‍

How long does it take to deploy AI insurance fraud detection in place?

For a first pattern on one line of business using existing systems, weeks rather than quarters. The initial work is connecting claims, policy, and provider systems and building the entity index; subsequent patterns reuse that foundation and deploy considerably faster.

Malavika Kumar
Director of Product Marketing
Published Sep 14, 2026