top of page

AI in Criminal Investigations Starts With Better Investigative Interviews

Oct 28, 2023
11 min read

Updated: Aug 31

Interviewing is data collection.


That sentence should change how a law enforcement agency or corporate investigative team thinks about every AI tool it is being sold. Transcription engines, case-linking software, large language models that summarize a file in seconds—none of them generate information. They process what investigators already collected. And the largest single source of information in most investigations is not a database. It is a person talking in a room to someone who is deciding what to ask next.

Two investigators review AI on a laptop during a criminal investigation in a glass-walled office while two men meet in another room.
Two investigators review AI output on a laptop during a criminal investigation. Better interviewing means better case data.

So the question underneath every AI purchase in criminal investigations is not "how good is the software." It is "how good is the interview?"


What AI actually does with an interview

The technology is less mysterious than the marketing suggests, and knowing what it does makes its dependence on input quality obvious.


Automatic transcription converts recorded interviews and body-worn camera audio into searchable text. Natural language processing reads that text for entities, relationships, and patterns—names, places, times, and how they connect. Large language models summarize, cross-reference, and surface links across cases that no detective has time to read end to end.


Every one of those functions operates on statements. Detail, specificity, and accuracy are the raw material. An algorithm can identify a pattern across ten thousand transcripts. It cannot tell you that a witness said what they said because the interviewer suggested it.

That is the entire problem, summed up in one sentence.


Four ways a bad interview poisons the dataset

Legacy accusatorial interviewing does not just produce a weaker interview; it produces a weaker dataset. It produces contaminated data, and contamination is worse than absence. A missing detail is a gap you can see. A wrong detail looks like a real one.


Leading questions plant information. When an interviewer supplies the detail and the subject adopts it, the transcript now contains a fact the person never independently knew. Nothing downstream can distinguish it from a genuine recollection.


Accusatorial pressure narrows the account. Confrontation triggers defensiveness. Defensive people say less, and what they say is shaped by the accusation rather than by memory. The field research is direct about this: accusatorial tactics are associated with more counter-interrogation behavior from suspects and less cooperation (Russano et al., 2024).


False confessions enter the record as fact. In a meta-analytic review of accusatorial and information-gathering methods, Meissner et al. (2014) found that accusatorial approaches increase the likelihood of eliciting confessions from innocent people, while information-gathering methods reduce it. A false confession is not a data-quality inconvenience. It is a corruption at the center of the case file, and every tool that touches that file inherits it.


Non-validated techniques generate confident noise. Interviewers trained to read micro-expressions, body language, or anxiety cues are not detecting deception; they are generating unreliable judgments and then acting on them. We have written at length about why that training makes people more confident without making them more accurate. Feed those judgments into a system and you have automated a bias.


Each of these was a problem before AI. AI makes them faster, and it makes them look authoritative.


What science-based interviewing puts into the system

Science-based interviewing is not a softer approach. It is a more productive one, and the productivity is measurable.


Free narrative and cognitive interviewing techniques increase the volume of usable detail. Memon et al. (2010) reviewed 25 years of Cognitive Interview research — 46 studies, 59 effect sizes, and nearly 2,900 participants — and found a large increase in correct details recalled compared with standard interviews. Errors rose only slightly and overall accuracy held roughly steady. More information, at the same reliability. That is precisely what a downstream analytic tool needs.


Rapport is a measurable interviewing skill, not a personality trait. The ORBIT framework—Observing Rapport-Based Interpersonal Techniques—scores interviewer behavior on motivational interviewing skills (acceptance, empathy, adaptation, evocation, and autonomy) and on whether interpersonal style is adaptive or maladaptive. Across 181 interviews with 49 convicted terrorists, more than 650 hours of tape, MI-consistent skill predicted more adaptive interviewing and significantly fewer counter-interrogation tactics: less passive resistance, fewer no-comment responses, and less deflection onto irrelevant topics (Alison et al., 2014). Rapport and active listening are how that skill shows up in the room. Cooperation is not a courtesy. It is the mechanism through which information reaches the case file.


Rapport is not befriending, sympathizing, or agreeing. It is a working engagement built on dignity, respect, and unconditional positive regard.

Strategic use of evidence produces statements you can test. Planning your questioning strategy through the evidence framing matrix (EFM) and holding evidence and disclosing it, if necessary or at all, deliberately lets the account and the evidence be compared rather than merged. In a controlled study, three hours of SUE instruction raised interviewers' accuracy from 56.1% to 85.4% (Hartwig et al., 2006). The skill is teachable—and being untrained sits at roughly a coin flip.


Holding evidence and disclosing it, if necessary or at all, deliberately lets the account and the evidence be compared rather than merged. In the foundational study, observers working from late-disclosure interviews reached 61.7% accuracy against 42.9% for early disclosure (Hartwig et al., 2005). For AI purposes, the value is not the accuracy figure—it is that the interview yields specific, checkable claims rather than a subject who already knows what you have.


The output of a science-based interview is a detailed, corroborable account that can be cross-referenced against records, digital evidence, physical evidence, and other statements. That is a dataset. The output of a confession-driven interview is "data" from pseudoscientific lie detection and often a short admission with very little in it, and this is the part agencies underestimate.


The same witness, two interviews

Abstract arguments about data quality are more effective when supported by a transcript. Here is the same witness to the same convenience-store robbery, interviewed two ways.


Interview A — closed, leading, confirmation-seeking.

Q: You saw the guy in the hoodie run out, right? A: Yeah, I think so. Q: Dark hoodie? Black or navy? A: Black, probably. Q: And he got into a car that was waiting? A: I guess there was a car. Q: Did he seem nervous? A: Yeah, he seemed nervous.

Four facts now sit in the record: a hoodie, its color, a waiting vehicle, and the suspect's emotional state. The interviewer implied and supplied every one of them. The witness supplied an agreement. If the hoodie was grey and there was no waiting car, nothing in this transcript will ever reveal that—and "seemed nervous" is not an observation at all, it is an inference the witness was invited to make.


Interview B — free recall, then depth.

Q: Tell me everything you remember, from the beginning, in as much detail as you can. Take your time. A: I was at the cooler. I heard the door chime and somebody was talking loud at the counter. When I turned around a guy was already walking out fast, pulling his hood up as he went. Grey, I think, with something printed on the back. He turned left out of the lot on foot. There was a car idling at the pump but I don't know if it was with him — it was still there when the police came.

One uninterrupted answer produces the position of the witness, the sequence of events, a corrected garment color, a printed graphic, a direction of travel on foot, and a vehicle explicitly flagged as unconnected. Every element is specific enough to check against the store's cameras, the pump transaction log, and the canvass.


That is the difference in a single exchange. Interview A yields four unverifiable assertions the interviewer authored. Interview B yields six checkable details the witness authored—plus an honest boundary around what they do not know.


Prosecutors know this before they read a word. Look at the two transcripts as shapes on a page rather than as text. In Interview A, the interviewer's blocks are long and the witness's are short. In Interview B that is reversed. Reviewing supervisors and prosecutors read that pattern at a glance because it shows who was doing the work—and the research says the pattern usually runs the wrong way. Across 80 real police interviews, interviewers spoke 55.58% of the words, and in 58.75% of those interviews the interviewer talked more than the person being interviewed. Open-ended questions made up less than 1% of all questions asked, yet they produced answers averaging about 91 words against roughly 10 words for yes-or-no questions, and the free narratives requested in only 13.8% of interviews averaged 286 (Snook et al., 2012).


Now imagine both transcripts inside a case-linking system. Interview A matches other cases on "black hoodie" and "waiting vehicle," neither of which happened. Interview B matches a printed graphic and a direction of travel, both of which it did. The software cannot tell the difference. It will process the fabricated pattern with exactly the same confidence as the real one, and it will hand an analyst a lead built on the interviewer's assumptions.


What clean input actually looks like

If you want a practical standard for what should be in a transcript before any tool touches it, this is the short version:


  • An uninterrupted free narrative captured before any specific questioning

  • Open questions that invite an account rather than confirm a theory

  • Basis of knowledge attached to every material claim—did they see it, hear it, or hear about it

  • Specifics that can be corroborated independently: times, positions, sequences, identifiers, people, communication, actions

  • Explicit uncertainty preserved rather than smoothed away, because "I don't know" is data

  • No interviewer-supplied detail anywhere in the record. That is contamination


Six criteria. An interview that meets them produces material worth analyzing. An interview that fails them produces a file that looks complete and is not.

ORBIT offers a complementary standard, and it is unusual in scoring the output as well as the behavior. Its Interview Yield Assessment codes what the interview actually produced: capability, opportunity, motive, and PLAT—people, locations, actions, and times (Alison et al., 2022). That is a data schema. It is close to what an analyst or a case-linking system needs and a far better standard for judging an interview than whether a confession was obtained.


The field evidence, not just the lab

Most interviewing research is laboratory work, and that is a fair criticism to raise. It is also increasingly out of date.


Russano et al. (2024) studied 90 investigators in the field and found that science-based tactics were associated with greater suspect cooperation, which in turn predicted more information disclosure and, indirectly, more admissions and confessions. Accusatorial tactics ran the other way: more counter-interrogation behavior, less cooperation, less information.


Russano et al. (2026) replicated that pattern with 43 investigators across three local agencies, tracing the same path from science-based tactics to cooperation to information disclosure to admissions. IXI contributed to that research as a practitioner co-author, which is also why we are careful about what it does and does not prove: it is correlational field data across a modest number of investigators, not a randomized trial.


The same pattern appears in work by a different research team, in a different country, with entirely different suspects. Alison et al. (2014) coded 181 real interviews with 49 convicted terrorists—878 tapes, more than 650 hours of recorded interrogation—and found that rapport-based interpersonal skill predicted fewer counter-interrogation tactics: less silence, fewer no-comment responses, and less deflection. Alison et al. (2022) then applied the same ORBIT framework to 35 interviews with 25 convicted child sexual abuse suspects, tracing the chain from interviewer behavior to suspect behavior to interview yield. Neither study was a laboratory simulation. Both used recordings of real interviews in real cases with people convicted of serious crimes.


This is what convergent validation looks like. Separate research teams working on opposite sides of the country on two continents, using different coding frameworks, studying populations as different as domestic criminal suspects, convicted terrorists, and child sexual abuse offenders, and measuring different outcomes—cooperation and information disclosure in one program, counter-interrogation tactics and interview yield in the other—arrive at the same result. Rapport-based, information-gathering methods produce more usable information. Accusatorial and maladaptive behavior produces less. Set the Canadian field data on how interviews are actually conducted (Snook et al., 2012) alongside the meta-analytic review of true and false confessions (Meissner et al., 2014), and the finding holds across method, population, and border.


There is a familiar objection to all of this, and it should be addressed: that this research is done in universities on undergraduates and has nothing to say about a real interview room. The studies above were run on real interviews with real suspects in real cases by researchers who went and got the tapes. An investigator remains free to disagree with what the findings mean. What is no longer available is putting your head in the sand and claiming that the findings do not exist. What all of this shows is consistent, and it is the point of this article. The methods that treat an interview as data collection produce more case data.


This is not only a police problem

Everything above applies to corporate investigations and loss prevention, and the exposure there is arguably worse.


HR teams, corporate security, and compliance functions are being sold the same category of tools—AI notetakers in investigative interviews, automated summarization of witness statements, sentiment analysis, and platforms that draft the investigation report. Most of those deployments happen without anyone auditing the interview practice underneath them. Many workplace investigators have had no formal interview training at all or worse legacy accusatory methods.


The consequences arrive differently, but a summary generated from a leading, contaminated interview becomes the document that supports a termination. A tool that flags "deceptive language" in an employee statement encodes a pseudoscience your organization will have to defend in a deposition. And unlike a criminal case, there is rarely a suppression hearing to catch it—the first real test is litigation, eighteen months later.


If your organization is running workplace investigations, the sequence is the same as it is for a police agency: fix the interview methods, then automate around it.


Before you buy the tool

Three steps, in this order.

  1. Stop the training that generates bad data. Any curriculum built on behavioral lie detection or confession-driven confrontation is manufacturing the contamination you will later pay software to analyze.

  2. Train the method that produces clean information. Rapport, question architecture, free recall, active listening, and disciplined evidence handling — with practice, not just slides. We use AI on this side of the problem too, running IXI-built scenarios so investigators rehearse the method that produces good data in the first place.

  3. Then deploy the technology. Applied to reliable input, these tools genuinely help. Applied to contaminated input, they scale the error and lend it a machine's credibility.


We made the fuller version of this argument in AI Stinks: Why Science-Based Interviewing Must Come First. The short version: get the horse in front of the cart.


Conclusion

AI is not going to fix an interview problem. It has no way to know it has one.


Agencies and organizations that invest in interviewing first will get real value from these tools, because their files will contain information worth analyzing. The ones that buy the software and skip the training will get faster, more confident, better-formatted versions of whatever was already wrong.


Interviewing is data collection. Everything downstream depends on how well it was done.


AI in Criminal Investigations References

Alison, E., Humann, M., Alison, L., Tejeiro, R., Ratcliff, J., & Christiansen, P. (2022). Observing Rapport-Based Interpersonal Techniques (ORBIT) to generate useful information from child sexual abuse suspects. Investigative Interviewing: Research and Practice, 12(1), 22–39.


Alison, L., Alison, E., Noone, G., Elntib, S., Waring, S., & Christiansen, P. (2014). The efficacy of rapport-based techniques for minimizing counter-interrogation tactics amongst a field sample of terrorists. Psychology, Public Policy, and Law, 20(4), 421–430. https://doi.org/10.1037/law0000021


Hartwig, M., Granhag, P. A., Strömwall, L. A., & Vrij, A. (2005). Detecting deception via strategic disclosure of evidence. Law and Human Behavior, 29(4), 469–484. https://doi.org/10.1007/s10979-005-5521-x


Meissner, C. A., Redlich, A. D., Michael, S. W., Evans, J. R., Camilletti, C. R., Bhatt, S., & Brandon, S. (2014). Accusatorial and information-gathering interrogation methods and their effects on true and false confessions: A meta-analytic review. Journal of Experimental Criminology, 10(4), 459–486. https://doi.org/10.1007/s11292-014-9207-6


Memon, A., Meissner, C. A., & Fraser, J. (2010). The Cognitive Interview: A meta-analytic review and study space analysis of the past 25 years. Psychology, Public Policy, and Law, 16(4), 340–372. https://doi.org/10.1037/a0020518


Russano, M. B., Meissner, C. A., Atkinson, D. J., Brandon, S. E., Wells, S., Kleinman, S. M., Ray, D. G., & Jones, N. J. (2024). Evaluating the effectiveness of a 5-day training on science-based methods of interrogation. Psychology, Public Policy, and Law, 30(2), 105–120. https://doi.org/10.1037/law0000422


Russano, M. B., Meissner, C. A., Jones, N. J., Rothweiler, J. N., Taylor, S., Cory, C., & Brandon, S. E. (2026). Evaluating the effectiveness of a practitioner-designed science-based interviewing and interrogation course. Legal and Criminological Psychology, 31, 273–300. https://doi.org/10.1111/lcrp.70021

Comments


bottom of page