The Tenth Man in the Room: What Four Centuries of Red Teaming Teach Investigators and Leaders
- Christian Cory

- Jun 18
- 18 min read
Updated: 2 days ago
In the International Spy Museum in Washington, D.C., a wall introduces visitors to the "Decision Room." The text makes a quiet but unsettling point: spies collect and analyze information, but deciding what to do with it falls to "consumers" of intelligence—heads of state, military commanders, diplomats, mayors. And the intelligence they receive "may come with recommendations or warnings, yet rarely is there 100% certainty." The case the museum uses to illustrate the problem is the 2011 operation against Osama bin Laden. Around the corner, a screen flickers with an invitation: Join the Red Team. So ask yourself before the verdict, before the headline, before the wrongful conviction with your name on the file: are you red teaming your cases—or are you one unchallenged assumption away from becoming the next cautionary tale?

That phrase—red team—names one of the oldest and most durable ideas in the history of organized decision-making. A red team is a group deliberately tasked with attacking its own side's plan, assumption, or conclusion (Development, Concepts and Doctrine Centre [DCDC], 2021). Its job is not to be right. Its job is to make sure the rest of the organization is not comfortably, catastrophically wrong. The label now stretches from military planning to cybersecurity, where a "red team" is the crew hired to breach your systems before an adversary does (Clark, 2014), but the underlying logic never changes. For investigators and the leaders who direct investigative work, red teaming is not just a military curiosity. It is a method and discipline for surviving the single most dangerous moment in any inquiry: the moment everyone agrees.
This article briefly examines four institutions that learned, often through difficult experience, to incorporate an adversary into their own thinking: the Prussian war academy, the Catholic Church, the Israeli intelligence community, and the modern CIA. Each found a different mechanism for the same problem. Together they sketch a practical playbook for anyone whose job is to reach the truth before the truth reaches them.
The problem red teaming solves
Start with the failure mode, because red teaming only makes sense as an answer to it.
Human reasoning under pressure tends to converge. We anchor to a first impression and adjust too little. We seek information that confirms what we already believe and discount what contradicts it. In groups, these private biases compound into something worse: the pressure to agree becomes its force, dissent feels disloyal, and the absence of objection gets mistaken for the presence of proof. Investigators know this pattern intimately—it is how a promising early suspect becomes the only suspect, how a working theory hardens into the theory, and how exculpatory evidence is quietly reinterpreted to fit a conclusion that was already reached.
it becomes more important than ever to have every perspective on the table.
The insidious part is that none of this feels like an error while it is happening. It feels like clarity. A team that has stopped questioning its own theory experiences that state as confidence, or even consensus. This is the premise behind formal red teaming doctrine: people and organizations invite failure in predictable ways, by degrees and almost imperceptibly, according to mindsets and biases formed by their own culture and experience (University of Foreign Military and Cultural Studies [UFMCS], 2016). Red teaming is the structural answer to a structural problem. You cannot reliably talk yourself out of a bias you cannot feel. That is why the advice to get rid of your biases is so unhelpful. However, you can be purposeful about it. Assign someone, formally and with cover, to argue the other side—and you can build that role into the process so it does not depend on any individual's courage.
Prussia: institutionalizing the adversary
The first modern attempt to make adversarial thinking routine came not from a philosopher but from a soldier with a sand table. In 1824, a Prussian artillery officer, First Lieutenant Georg Heinrich Rudolf Johann von Reisswitz, refined a game his father had developed and presented it to the Prussian general staff. They were impressed enough that the chief of staff reportedly declared, "This is not a game! This is training for war!" The king ordered a set sent to every regiment (Kriegsspiel, n.d.).

Kriegsspiel—literally "war game"—was built on two innovations that matter as much to an investigator as to a general. The first was the umpire. Two players commanded opposing forces, but a third participant, the Vertraute or "trusted one," served as an impartial judge: he adjudicated outcomes, applied the rules, and—crucially—controlled what each side was allowed to know (Kriegsspiel, n.d.). The second innovation followed from the first: fog of war. Because the umpire withheld information, each commander had to act on incomplete and sometimes false pictures of the battlefield, exactly as they would in reality.
What made Kriegsspiel revolutionary was not the dice or the maps. It was the institutional decision to make a living, thinking opponent a permanent fixture in how Prussian officers learned to plan. You did not test your plan against your optimism. You tested it against a human being whose entire assigned purpose was to defeat you, refereed by a third party with no stake in your being right. When Prussia routed France in 1870, many credited the army's wargaming tradition for its superior planning (Kriegsspiel, n.d.). The lasting lesson is procedural rather than military: adversarial pressure works best when it is built into the system, scheduled and staffed, rather than summoned in a crisis. By the time you feel you need a red team, you are usually too committed to have an honest one.
The Church: dissent as office, not attitude
Two and a half centuries before Reisswitz, the Catholic Church had already turned skepticism into a salaried position. In 1587, under Pope Sixtus V, the Church formalized the role of the Promotor Fidei—the Promoter of the Faith—universally known by its nickname, the advocatus diaboli, the Devil's Advocate (NPR, 2013). When a candidate was proposed for sainthood, this canon lawyer was appointed to argue against the canonization: to take a deliberately skeptical view of the candidate's character, to hunt for holes in the evidence, and to argue that the miracles attributed to the candidate were exaggerated or fraudulent (NPR, 2013).
This is one of history's cleanest expressions of an idea investigators should internalize. The Church did not wait for a skeptic to volunteer, and it did not treat skepticism as a personality flaw to be tolerated. It created an office. Someone was required to build the strongest possible case against the proposition the institution was emotionally and politically inclined to approve. The burden of proof ran uphill against the preferred conclusion, and a specific, named person was responsible for keeping it there.

In 1983, Pope John Paul II reformed the canonization process and effectively sidelined the Devil's Advocate, in part because the role had become a bottleneck (NPR, 2013). The result was striking: John Paul II went on to canonize nearly 500 people and beatify more than 1,300—against fewer than 100 canonizations by all of his twentieth-century predecessors combined (NPR, 2013). One can read that as welcome efficiency or as a cautionary tale, but the mechanism is unambiguous. Remove the institutional skeptic and the throughput of approvals soars. For an investigative unit, that is precisely the warning. When you dismantle or marginalize the function whose job is to say "not yet, not proven," cases close faster—but this does not necessarily mean they are more sound. Speed and rigor are not the same virtue, and red teaming is one of the few practices that protects the latter from the former.
Israel: the catastrophe that created a contrarian unit
If Prussia and the Church show red teaming working by design, Israel shows what it costs to learn the lesson by disaster.
In the autumn of 1973, Israeli military intelligence was operating under what insiders called ha-Konseptzia—"the Concept." The Concept held that Egypt would not go to war until it could field aircraft capable of striking deep into Israel, and that Syria would not attack without Egypt. From that premise, an avalanche of warning signs was explained away. Egyptian and Syrian forces massing on the borders were read as exercises. The intelligence was not missing; the interpretation was wrong. As U.S. Secretary of State Henry Kissinger later put it, "There was no lack of intelligence; it was the interpretation to the reports that was faulty" (GlobalSecurity, n.d.). On October 6, 1973—Yom Kippur—Egypt and Syria attacked simultaneously, nearly overrunning Israel before it recovered.

The post-war Agranat Commission dissected the failure and reached a conclusion that should be tattooed on the wall of every analytic shop: the disaster was not primarily a failure to collect but a failure to challenge a dominant assumption. In response, Israeli Military Intelligence created a small, permanent unit dedicated to institutionalized dissent—now known as the Devil's Advocate Unit, or Ipcha Mistabra (Devil's Advocate Unit, n.d.). The name is Aramaic, drawn from a Talmudic argument, and translates roughly as "the opposite is reasonable" or "on the contrary, it appears the reverse" (Devil's Advocate Unit, n.d.). The unit's mandate is to examine unlikely scenarios, attack the prevailing consensus of the Research Division, and write the alternative assessment that nobody else wants to write (Devil's Advocate Unit, n.d.).
In popular culture the concept hardened into the so-called "Tenth Man" rule, dramatized in the film World War Z: if nine people in a room examine the same evidence and reach the same conclusion, the tenth is obligated to assume they are all wrong and to make the case that they are. The cinematic version is tidier than the historical reality, but the underlying doctrine is real and consequential. Its value is not theoretical. In September 2023, the head of the Devil's Advocate Unit reportedly issued repeated explicit warnings over a three-week span about the potential for a large-scale Hamas assault, challenging the consensus that Hamas lacked the intent and capability to escalate (Devil's Advocate Unit, n.d.). Acting on those warnings would have strengthened the case for the unit. It strengthens it—and it adds a second lesson on top of the first. Standing up a red team is necessary but not sufficient. Leadership has to be willing to actually hear it, which is a discipline of its own.
Abbottabad: red teaming under radical uncertainty
Which brings us back to the Spy Museum's Decision Room and the operation it chose to dramatize.
By early 2011, the CIA had assembled a circumstantial case that a high-value individual living behind high walls in Abbottabad, Pakistan, was Osama bin Laden. No one had seen his face. There was no photograph, no intercepted voice, no DNA—only a pattern of behavior, a courier network, and a compound built for secrecy. The analysts who had lived inside the hunt were confident. Others were not. And the agency did something that is the entire point of this essay: it deliberately attacked its own conclusion before asking the President to act on it.
The CIA convened red teams—analysts who had not worked the case—to independently review the same evidence and reach their own judgments, precisely so the leadership would not be relying on a single, possibly self-reinforcing line of analysis. One administration official summarized the approach plainly: "We conducted red-team exercises and other forms of alternative analysis to check our work. No other candidate fit the bill as well as bin Laden did" (Killing of Osama bin Laden, n.d.).

The result of that exercise is the most honest thing about it. The estimates did not converge on a comforting number. The lead analyst who had spent years on the hunt put his confidence at 90 to 95 percent. The deputy director of the CIA landed around 60 percent. Members of the red team came in markedly lower—some near 40 percent—and across the deliberations the President was presented with a spread running anywhere from roughly 30 to 95 percent that the man in the compound was bin Laden (NPR, 2014). President Obama's own synthesis was famously deflationary: he concluded it was, in effect, a coin flip—"50-50"—and decided to launch the raid anyway (NPR, 2014).
There are two lessons stacked here, and investigators should separate them carefully. The first is that red teaming does not exist to produce certainty; it exists to produce an honest map of uncertainty. The value of the exercise was not that it told Obama bin Laden was there. It was that it surfaced, in plain view, how thin and contested the case actually was—so that the decision could be made with eyes open rather than on a wave of manufactured consensus. The second lesson is about the relationship between analysis and decision. The red team's job was to sharpen the estimate; the leader's job was to decide and own the consequences. Obama did not wait for the doubt to disappear, because it never would. He acted in full view of it. Good red teaming and decisive leadership are not in tension. The first is what makes the second responsible rather than reckless.
Red teaming in practice: notes from the field
The four cases above are institutional, drawn from armies and churches and intelligence services. The following are personal—drawn from years in investigations and crisis work—but they are the same idea wearing a badge. The mechanism never changes. Someone has to be assigned to push back, and the room has to be built so that pushing back is safe.
In crisis and hostage negotiation, we used something we called the red negotiator. This was usually a more senior negotiator on scene. When we began passing information from the negotiation operations center—the NOC—to brief the on-scene commander, we would lay out the situation as we understood it from our end, our assessment, and our recommendation. That was the moment the red negotiator stepped in and played the contrarian. Seasoned negotiators have seen a great deal across a great many scenes—barricades, suicidal subjects, co-located hostage situations, genuine hostage takings—and the red negotiator's job was to make sure we were not taking shortcuts in our own thinking. We wanted to be certain our assessment was not simply what we expected to happen next, because no two situations are ever the same. Sometimes the recommendation stayed exactly the same, but we added caveats so the commander understood the limits of what we were telling him. Other times we adjusted the recommendation outright based on what the contrarian had surfaced. Either way, the commander received a sharper, more honest picture—not a smoother one.
As you grow as a leader, one thing is reliably true: things will pass you by. Your perspective changes the further you move up the chain of command and the more distance opens between you and the rank and file. At the same time, there has never been a period in history with as much digital and physical evidence available to investigators—license plate readers, phone extractions, tower dumps, an entire city's worth of cameras running from government systems down to doorbell footage. There is more evidence out there to help solve a case than ever before, and that abundance is exactly why red teaming matters now. It lets you challenge your assumptions about the evidence and force alternative explanations for the gaps, and it keeps the investigation thinking. We have all seen the investigative disasters that end up as streaming documentaries, where you watch from the outside and ask, how could they have missed that? Red teaming is how you avoid being that case. You build an environment in which the most junior detective can challenge the most senior detective, because they all hold perspectives worth hearing. And as leaders we have to remember that sometimes it is best for us to speak last. If the person with the most authority states a conclusion first, the room tends to fall in line behind it—the bandwagon forms, and the people who might have objected decide it is not worth it (Janis, 1972; TruThinking Corp, 2023). Avoiding that is not passive. We have to foster it deliberately, building the act of asking our people for their read into the process rather than hoping it happens on its own.
One story stays with me. We would hold what we called felony reviews when we had a high-profile case or a homicide, presenting it in front of a group of attorneys. I was watching a new homicide detective present his case. It went well; the case got charged. There was the usual back-and-forth among the attorneys that anyone who has been in that room comes to expect. On the way out, the new detective pulled me aside and asked, in so many words, why one of the attorneys had to be such a jerk. I told him: because that man is the smartest guy in the room. He is playing the defense attorney. He is telling the other attorneys exactly how he intends to come after them in court, which is why you walked out with a short list of homework and items to check off—so they are ready when the motions start coming in and trial preparation begins. He had just made the detective's case stronger and far less vulnerable to attack. That is a red team doing its job, and it is worth recognizing as a gift rather than resenting it as an obstacle.
There is a particular trap I warn every class about, because it is the engine behind so many of those cascading errors: techniques that sound like science and do nothing but confirm what we already believe. Neuro-linguistic programming and its eye-movement "tells," microexpressions, baselining, and the Behavior Analysis Interview—they carry the vocabulary of rigor, and that is exactly what makes them dangerous. They do not make our cases more accurate. They make us more confident, which in an investigation is the worse of the two problems. When a method tells you the calm suspect is "too controlled" and the nervous one is "showing deception," it cannot be wrong, because every outcome confirms it. A method that no result can disprove is not science; it is a mirror.
The research is not close. People reading demeanor to separate lies from truth land around 54 percent accuracy—barely above a coin flip—and trained officers do no better than anyone else (Bond & DePaulo, 2006). The NLP eye-movement claim was tested head-on and failed three separate ways (Wiseman et al., 2012). The Behavior Analysis Interview does not reliably distinguish liars from truth-tellers and actively risks branding the innocent as deceptive (Vrij, Mann, & Fisher, 2006). And microexpressions—the crown jewel of the genre—turn out to be rare and to appear about as often in honest people as in liars (Porter & ten Brinke, 2008). It is the equivalent of telling you I scored 100 percent on the test, then admitting I only counted the questions I got right. The people who sell this will concede that "nothing is 100 percent," but they will never hand you an independent, peer-reviewed number for how accurate their method actually is—because outside their own materials, it does not survive testing. These are bias-bolstering tools, and that is all they are. They belong nowhere near the interview room, and yet that is precisely where they get taught. Remember who carries the risk: the instructor who certified you will never sit in your witness chair to defend what they taught. You will. If you want an approach that earns its confidence, that lane already exists—Science-Based Interviewing (SBI), built for American jurisprudence and law enforcement, seeks reliable information through rapport, sound questioning, and the Strategic Use of Evidence (SUE) rather than reading twitches (Meissner et al., 2014).
Finally, there is the matter of false confessions, which are among the worst outcomes in any miscarriage of justice, especially when paired with a wrongful conviction. When I speak to classes and audiences, I point out that a great deal comes before a false confession—biases, tunnel vision, and false case information fed to a suspect through legacy interviewing techniques and pseudoscientific lie detection (Kassin, 2005). What you usually see is not a single mistake but several cascading errors that snowball through an investigation (Leo & Davis, 2010). I have watched it happen when a team becomes fixated on a case theory—because it fits neatly or because it was someone's idea and being right earns them a gold star for the day. Tunnel vision does not always end in a false confession or a wrongful conviction (Findley & Scott, 2006). Sometimes it just produces arguments among team members, or wasted resources, or squandered hours during the precious early window of an investigation. Red teaming keeps your team thinking. As a leader, you create an environment where everyone can speak up, offer alternative explanations, and play devil's advocate (aka a defense attorney) against the leading theory before the interviews so you know exactly what you need to clear up. As I always say: rule in and rule out. This instinct is increasingly being formalized; modern quality-control frameworks for fraud and corruption investigations build adversarial review directly into the process as a safeguard against exactly these cascading errors (Willems et al., 2025). And frankly, as technology keeps growing and you find yourself leading a new generation of detectives who investigate in ways you never did, it becomes more important than ever to have every perspective on the table.
From history to practice: building dissent into investigations
Four institutions, four mechanisms—an umpire, an office, a unit, an exercise—and a single recurring insight: you cannot rely on a team to spontaneously challenge a conclusion it is invested in. You have to engineer the challenge. For investigators and the leaders who run investigative organizations, the history translates into a handful of concrete practices—several of which modern militaries and defense organizations have already codified into working doctrine (DCDC, 2010, 2021; UFMCS, 2016; Defense Science Board [DSB], 2003).
Assign the role; do not hope for the trait. The Church did not wait for a brave skeptic, and neither should you. On any significant case, name a specific person—rotating the duty so it does not become one individual's reputation—whose explicit, sanctioned job is to build the strongest case against the leading theory. Give them cover. The whole point of an office of dissent is that disagreement becomes a duty performed rather than an act of insubordination risked.
Red team the theory of the case before you commit, not after it collapses. Prussia built the adversary into training precisely so plans met resistance while they were still changeable. The natural moment to convene a red team in an investigation is before an arrest, before a charging decision, before a final report goes up the chain—at exactly the point when the team feels most confident and is therefore most exposed. A red team called only after something has gone wrong is a post-mortem, not a safeguard.
Hunt for the assumption, not just the evidence. The Yom Kippur failure was not a gap in collection; it was a Concept that quietly converted every new fact into confirmation. Ask of any investigation: what is our equivalent of the Concept? What single load-bearing assumption, if it were false, would make our whole picture collapse—and what would the evidence look like if it were false? Often the most important question is not "what supports our theory" but "what would we expect to see if our theory were wrong, and do we see it?" Structured question sets built for exactly this purpose—asking what we assume, what we know, and what we would have to believe for the opposite to be true—turn that instinct into a repeatable drill (TruThinking Corp, 2024).
Separate the estimate from the decision. The Abbottabad case works because analysis and command were kept distinct. The red team's product was a calibrated, contested probability—not a recommendation to act. Investigators serve leaders better by reporting genuine uncertainty in honest ranges than by laundering a hunch into false confidence to make a decision easier. A "60 percent, and here is exactly why the other 40 is not crazy" is worth more than a "we're sure."
Protect the contrarian, then actually listen. The Israeli unit's 2023 warnings are a reminder that a red team is worthless if its conclusions are filed and ignored. The hardest part of this discipline is not standing up to the function; it is the leadership maturity to invite a conclusion you do not want, sit with the discomfort it creates, and weigh it on the merits. A red team that is tolerated but never heeded is theater. A red team that can change a decision is a safeguard.
The discipline of being your own adversary
The common thread from Reisswitz's sand table to the Situation Room is not pessimism. It is humility operationalized. Each of these institutions confronted the same human tendency—the drift toward agreement, the comfort of a settled story—and refused to trust itself to resist it through good intentions alone. They built the doubt into the architecture: an umpire who controls the truth, an advocate paid to attack the saint, a unit whose loyalty is expressed through dissent, an exercise that turns its analysts against their own best case.
For an investigator, the most dangerous moment is never the one where the file is thin and everyone is arguing. It is the one where the file feels complete and everyone agrees. That is the moment the tenth man exists for. The work of building better thinking, better leadership, and better investigations comes down to a single uncomfortable habit: before you decide you are right, assign someone—maybe yourself—to prove that you are wrong, and mean it. The Spy Museum's invitation is the right one to accept. Join the red team. Especially when the red team is aimed at you.
How to Red Team
Red teaming isn't a personality trait or a meeting agenda item. It's a discipline—one that has to be built, practiced, and protected from the pressure that closes it down. The teams that do it well don't have smarter people. They have a structure that makes it safe to be wrong out loud, early, while it still costs nothing.
If you're interested in improving your investigations, your team's thinking, or your leadership, IXI runs on line, in person, and hybrid workshops on exactly this. cory@ixi.solutions
Red Teaming References
Bond, C. F., Jr., & DePaulo, B. M. (2006). Accuracy of deception judgments. Personality and Social Psychology Review, 10(3), 214–234.
Clark, B. (2014). RTFM: Red team field manual. CreateSpace Independent Publishing Platform.
Defense Science Board. (2003). Report of the Defense Science Board Task Force on the role and status of DoD red teaming activities. U.S. Department of Defense, Office of the Under Secretary of Defense for Acquisition, Technology, and Logistics.
Development, Concepts and Doctrine Centre. (2010). A guide to red teaming (DCDC guidance note). UK Ministry of Defence.
Development, Concepts and Doctrine Centre. (2021). Red teaming handbook (3rd ed.). UK Ministry of Defense.
Devil's Advocate Unit. (n.d.). In Wikipedia. Retrieved June 17, 2026, from
Findley, K. A., & Scott, M. S. (2006). The multiple dimensions of tunnel vision in criminal cases. Wisconsin Law Review, 2006(2), 291–397.
GlobalSecurity. (n.d.). Yom Kippur War: Grand deception or intelligence blunder. GlobalSecurity.org. Retrieved June 17, 2026, from
Janis, I. L. (1972). Victims of groupthink: A psychological study of foreign-policy decisions and fiascoes. Houghton Mifflin.
Kassin, S. M. (2005). On the psychology of confessions: Does innocence put innocents at risk? American Psychologist, 60(3), 215–228.
Killing of Osama bin Laden. (n.d.). In Wikipedia. Retrieved June 17, 2026, from
Kriegsspiel. (n.d.). In Wikipedia. Retrieved June 17, 2026, from
Leo, R. A., & Davis, D. (2010). From false confession to wrongful conviction: Seven psychological processes. Journal of Psychiatry & Law, 38(1–2), 9–56.
Meissner, C. A., Redlich, A. D., Michael, S. W., Evans, J. R., Camilletti, C. R., Bhatt, S., & Brandon, S. (2014). Accusatorial and information-gathering interrogation methods and their effects on true and false confessions: A meta-analytic review. Journal of Experimental Criminology, 10(4), 459–486.
NPR. (2013, March 3). Who is the 'Devil's Advocate'? National Public Radio.
Porter, S., & ten Brinke, L. (2008). Reading between the lies: Identifying concealed and falsified emotions in universal facial expressions. Psychological Science, 19(5), 508–514.
NPR. (2014, July 23). In facing national security dilemmas, CIA puts probabilities into words. National Public Radio.
TruThinking Corp. (2023). Overcoming cognitive bias and groupthink (Red Team Coaching, Module 2) [Workbook]. TruThinking Corp.
TruThinking Corp. (2024). Think-write-share & six strategic questions (Red Team Coaching, Module 1) [Workbook]. TruThinking Corp.
University of Foreign Military and Cultural Studies. (2016). The applied critical thinking handbook (Version 8.1). U.S. Army.
Vrij, A., Mann, S., & Fisher, R. P. (2006). An empirical test of the behaviour analysis interview. Law and Human Behavior, 30(3), 329–345.
Willems, T., Stahn, C., Frey, D., & Angotti, A. (Eds.). (2025). Quality control in fraud and corruption investigations. Torkel Opsahl Academic EPublisher.
Wiseman, R., Watt, C., ten Brinke, L., Porter, S., Couper, S.-L., & Rankin, C. (2012). The eyes don't have it: Lie detection and neuro-linguistic programming. PLOS ONE, 7(7), e40259.



Comments