Introduction

Abraham Wald, born in 1902 in the city of Cluj, which at the time was part of Austria-Hungary, was a distinguished mathematician whose work revolutionized several fields within statistics and decision theory. Wald was Jewish and had to flee the Nazi advance across Europe, which led him to move to the United States in 1938. There, he became a member of the Statistical Research Group (SRG) at Columbia University, where he worked throughout World War II.


At one point during the war, the Allies were losing a large number of aircraft in combat, and faced the need to reinforce them to improve their odds of survival. To do this, they drew up diagrams of the planes that returned from missions and analyzed the areas that had been hit by enemy fire. This led them to conclude that they should reinforce the areas showing the most damage, since those appeared to be the most vulnerable.


Survivorship Bias Impacted Areas

However, since they weren't getting the results they expected, they asked Abraham Wald to review this analysis. After looking at the diagrams, Wald immediately noticed that these planes, despite being riddled with bullet holes, were in fact the ones making it back to base, and that the real question they should be asking was: what happened to the planes that didn't come back? Why weren't they in the analysis? Those missing planes had clearly been hit in more critical areas than the ones that returned, and therefore those were the areas that actually needed reinforcing.


Survivorship Bias Reinforced Areas

Wald's approach led to reinforcing the areas opposite the ones that were hit


Wald's approach, grounded in the data that had been ignored, was revolutionary and helped save countless Allied lives, and it gave rise to the concept of Survivorship Bias.


What Is Survivorship Bias?

Survivorship Bias, then, is an "error" that happens when we only account for the "survivors" of a process or event, ignoring those who didn't make it through. Just as in the analysis of the planes that survived, this concept refers to the tendency analysts have to focus only on what's visible or successful, which leads to incomplete and inaccurate conclusions.


There are many examples of this bias in everyday and professional life that can lead us to incorrect conclusions by relying on incomplete or skewed data. For example, if a company only looks at employees who've been promoted, and not those who were let go, it might conclude that the promotion resulted from a certain behavior, when in reality other factors may have been involved.


In finance, only analyzing companies or stocks that succeeded in the market, and not the ones that failed or lost value, can also lead to bad investments.


In medicine, only studying patients who responded to a treatment, and not those who didn't, can lead to prescribing ineffective medications.


In e-commerce, we might analyze the number of shopping carts that were completed, and not the ones that were abandoned, which would lead us to the wrong conclusions about the purchase process. If we focus only on the 100 carts that succeeded, we're ignoring the fact that 300 never went through. This might lead us to think the purchase process is working well, when in reality there's a 75% abandonment rate. Understanding the "why" behind that 75% is vital to figuring out whether there's a problem in the purchase process, and how to fix it.


We could list many examples like this, where Survivorship Bias leads us to the wrong conclusions. That's why it's important to keep this concept in mind when making decisions, since the data we don't see can be just as important as the data we do see. It's essential for avoiding faulty conclusions, and that's where Decision Theory comes in, a statistical approach that helps us make rational decisions under uncertainty.


Decision Theory

Decision Theory is concerned with analyzing how a person chooses the action that, among a set of possible options, leads to the best outcome, given their own preferences and goals. In other words, the theory seeks to identify the option that maximizes benefit or utility for the decision-maker, based on what they consider most valuable or desirable.


The theory proposes that decisions should be made using all the relevant information available. The problem arises, however, when the data is biased or incomplete, as is the case with Survivorship Bias.


To address this problem, the theory urges us to identify and consider information that's missing or not visible. This means actively asking "What data aren't we seeing?" or "What factors are we ignoring?" The key is making sure every possible variable, both obvious and hidden, gets included in the analysis before making a decision.


This comprehensive approach helps avoid wrong conclusions and incomplete decisions that can arise from relying solely on visible data or success stories. By considering the full picture and evaluating every relevant variable, decision-making becomes more precise and effective.


Fundamental Elements of Decision Theory


To understand how informed decisions get made, it's crucial to know the fundamental elements of Decision Theory. These elements form the basis for analyzing and selecting the best options under uncertainty:


  • Decision-Maker: the decision-maker can be a person, an organization, or a system that needs to make a decision. This agent aims to maximize a specific criterion, like utility, benefit, or efficiency. The decision-maker's role is fundamental, since it defines the goals and parameters that will guide the decision.
  • Set of Alternatives: every decision involves choosing among several possible options. Each alternative has different consequences, and Decision Theory helps evaluate which option is most favorable based on the established goals. This analysis lets you compare alternatives in terms of their expected outcomes and associated costs.
  • Outcomes or Consequences: decisions lead to one or more possible outcomes, which can be certain or uncertain. The theory provides tools for handling both cases. Evaluating consequences is key to understanding the impact of each alternative and making more informed decisions.
  • Decision-Maker's Preferences: decision-makers often have subjective preferences about outcomes. Decision Theory uses utility functions to quantify these preferences, letting you measure the satisfaction or value tied to each possible outcome. This helps align decisions with personal or organizational goals and values.
  • Probabilities: under uncertainty, it's essential to assign probabilities to the different possible outcomes. Statistical and probabilistic techniques, like Bayes' rule, are used to adjust probabilities as new information becomes available. This approach lets you make decisions based on a precise assessment of probabilities and associated risks.

These elements provide a solid framework for decision-making under uncertainty. By carefully weighing each aspect, decision-makers can improve the quality and effectiveness of their choices, minimizing the impact of biases like survivorship bias and maximizing the odds of reaching their desired goals.


Types of Decisions

Decision Theory covers several types of decisions, classified mainly by the certainty or uncertainty of the outcome and the nature of the choice:


Decisions Under Certainty: in this type of decision, the outcome of each action is known with certainty. The decision is based on complete, accurate information, with no uncertainty about future outcomes. Deciding on the most efficient method for a task where all the factors and outcomes are well known is a good example.


Decisions Under Uncertainty: in these decisions, the outcome of each action isn't fully known. The uncertainty can be total or partial, and probability theory tools are used to manage it.


Decisions Under Risk: the probability distribution of possible outcomes is known. Decisions are made based on the expected value of each option. This could mean, for example, deciding to invest in a stock knowing the probabilities of different future returns.


Intertemporal Decisions: these decisions involve choosing between options that have consequences across different time periods. It's about weighing the present value of future benefits against current costs. An example would be choosing between spending money now on a luxury or investing it for future benefits.


Social Decisions: in this case, decisions are made collectively or within an organizational structure. This means considering not just individual interests, but also the impact on the group or organization. Examples include deciding how to allocate resources in a company or planning public or social policy.


Decisions Under Conditions of Complexity: these decisions involve a high degree of complexity in calculating expectations or within the organization making the decision. The theory focuses on understanding how difficult it is to find the optimal solution. An example would be deciding on a company's marketing strategy in a highly competitive, ever-changing market, like a soccer club.


Decisions Between Incommensurable Goods: this type of decision involves choosing between options that can't be measured using the same unit or criterion, like choosing between investing in different types of assets that have different qualitative and quantitative valuations.


The Paradox of Choice: this phenomenon occurs when having too many options leads to less satisfaction with the decision, or even analysis paralysis. Think of facing an overwhelming number of product options in a store, which can lead to a less satisfying decision, or to making no decision at all.


These types of decisions show the diversity of scenarios where the principles of the theory can be applied. Each one requires a specific approach and the right tools to ensure informed, effective decision-making.


Applications in Data

Decision Theory is fundamental in the field of data science and data analysis. When facing large volumes of information and situations of uncertainty, informed decisions need to be made that maximize the value of the data while minimizing the associated risks. Some key applications of the theory in data science include:


Marketing Campaign Optimization: in marketing, the theory is used to select the most effective strategy. Historical data from previous campaigns is analyzed to calculate the expected value of each strategy. For example, regression models can be used to predict the conversion rate of each strategy and select the one most likely to maximize return on investment (ROI).


Classification Models in Machine Learning: in financial fraud detection, the theory is applied to build predictive models that classify transactions as fraudulent or not. It helps adjust the probabilities of different outcomes (fraud or no fraud) and assign weights to relevant features. A decision tree model, for example, can evaluate a transaction's features and make decisions based on minimizing the cost of false positives and false negatives.


Investment Portfolio Management: in finance, it's used to select investments that maximize expected return while minimizing risk. Using Expected Utility Theory, you can calculate the probabilities of different investment returns and make decisions based on the balance between risk and reward. This can mean optimizing asset allocation within a portfolio to achieve the best risk-adjusted return.


Resource Planning in Projects: when planning projects with multiple tasks and limited resources, the theory helps allocate resources in a way that maximizes efficiency and meets project deadlines. Different resource allocation scenarios can be evaluated to select the one offering the best balance between cost, time and quality.


Supplier Selection in the Supply Chain: when choosing among several suppliers for a raw material, the theory helps evaluate and compare offers. This can involve assessing criteria like cost, quality, delivery time and reliability. Weights can be assigned to each criterion, and the decision that maximizes total value for the company can be selected.


Social Media Sentiment Analysis: in social media sentiment analysis, the theory helps interpret sentiment classification results to make decisions about communication or marketing strategies. Natural language processing (NLP) models can classify user comments and help adjust campaigns based on the prevailing sentiment.


There are many applications of Decision Theory in the field of data science and analysis, and integrating it into decision-making processes can significantly improve the quality and effectiveness of the choices made. By considering the theory's fundamental elements and applying them to situations of uncertainty, you can maximize the value of data and minimize associated risks, contributing to the success of projects and organizations.