This post is adapted from a webinar I gave for the SPE Data Science and Engineering Analytics Technical Section in December 2024. The recording is available on YouTube.

If you have been reading the news for the last forty years, you have probably seen at least one story about how a glass of red wine each day reduces the risk of heart disease. You have probably also seen at least one story about how any amount of alcohol consumption is bad for you. These stories are often presented as if the scientific community has changed its mind, or as if journalists are causing trouble by exaggerating small effects from boring papers. That does happen sometimes, but it is not the most interesting version of the problem.

The more interesting version is that the studies can be reasonable, the authors can be acting in good faith, the journalists can summarize them accurately, and the conclusions can still point in opposite directions. This is not because science is fake or statistics are fake or everyone involved is bad at their job. It is because the thing we want to know is often not the thing we measured.

The alcohol example is a good one because the direction of the confusion is easy to see after someone points it out. Moderate drinkers, at least in many western countries, are not just people who consume one or two alcoholic drinks per day. They are also more likely to be socially active, more likely to go out to dinner with friends, and may differ in lots of other ways that affect health outcomes. People who abstain entirely are not just people with zero drinks per day. Some abstain for religious reasons, some abstain because alcohol makes them feel bad immediately, some abstain because they previously had a problem with alcohol, and some abstain because they already have liver disease or another health condition that makes drinking obviously unwise.

If you compare those two groups directly, you are not only comparing alcohol consumption. You are comparing two populations that differ in a lot of other ways. It is therefore not surprising that the moderate drinkers sometimes look healthier than the abstainers. But if you could take the same moderate drinkers, remove the alcohol from their lives, and watch the alternate version of their health outcomes unfold over decades, you would probably get a different answer. If you could take the same abstainers, add alcohol, and watch that other future unfold, you would probably get a different answer again.

This is the basic problem causal inference is trying to solve. We want to compare two futures, but we only get to observe one.

The usual machine learning formulation does not solve this problem. A discriminative model, whether it is a linear regression, random forest, gradient boosted tree, or neural network, estimates something like the expected value of an outcome given a set of inputs. Given age, exercise, income, social network size, and alcohol consumption, what health outcome should we expect? That can be a useful question. It is not the same as asking what would happen if alcohol consumption changed while everything else about the person remained fixed.

The model does not care whether "having lots of friends" causes "exercising more", or whether both are downstream of some third thing, or whether the measurement is mostly proxying for socioeconomic status. If these variables help predict the outcome, the model will use them. In many machine learning applications this is exactly what we want. In causal applications it can be exactly the problem.

What we would like instead is a treatment effect. In the alcohol case, the treatment might be the number of standard drinks consumed per day. In an oil and gas example, the treatment might be the amount of proppant used in a completion. In a product experiment, it might be whether a user saw the new onboarding flow. In all of these cases the structure of the question is the same: what is the difference between the outcome under one intervention and the outcome under another intervention?

The cleanest way to estimate this is a randomized controlled trial. If treatment assignment is random, then the covariates we are worried about should balance out across groups as the sample size gets large. The people assigned to the treatment group and the people assigned to the control group should be similar in all the annoying, hard-to-measure ways, because nothing about who they are determined which group they entered. In the limit, the difference in outcomes between those groups converges on the causal effect of the treatment.

This is very nice when you can do it. It is also very often unavailable. It would be unethical to force half the planet to drink alcohol for fifty years so that we can measure cancer outcomes. It would be infeasible to give everyone a billion dollars and then measure changes in health. It is often commercially unappealing to run a perfect randomized business experiment when the bad arm could cost millions of dollars. It is sometimes operationally impossible to assign treatments randomly because the treatment has already happened, the wells have already been drilled, the customers have already churned, and the dataset is what it is.

The good news is that this has not ruined everything. The bad news is that we have to be careful.

One way to be careful is to make the comparison more local. If we believe some variable is influencing both treatment assignment and the outcome, we can stratify the data by that variable and estimate treatment effects inside each stratum. In the alcohol example, we might want to compare people with similar prior health conditions, similar social patterns, and similar propensity for addiction. In the classic Simpson's paradox version of the problem, aggregating across groups gives an effect in one direction while estimating the effect inside each group gives the opposite answer.

This family of methods shows up under several names: stratification, segmentation, matching, blocking back-door paths. The names differ because this problem has been rediscovered by several fields. Medicine, public health, economics, sociology, and machine learning all have versions of it. The core idea is pretty intuitive though. If the treated and untreated groups are different in ways that matter, find smaller groups where they are more comparable and make the comparison there.

Another way is to look for something outside the system that changes treatment assignment without directly changing the outcome. This is the natural experiment or instrumental-variable approach. In the alcohol example, we might imagine some exogenous shock that changes access to alcohol for reasons unrelated to anyone's health. In an oil and gas example, maybe there is a temporary supply constraint on sand. If the supply constraint changes proppant loading but does not change rock quality, takeaway capacity, or the other variables that directly determine production, then it can help us estimate the causal effect of proppant.

These examples are opportunistic. You do not get to schedule natural experiments on demand, and most things that look like instruments stop looking like instruments once you think about them for more than five minutes. The supply constraint cannot also be changing the number of crews available, the quality of the completion, the date the well comes online, or anything else that has its own path to production. Otherwise the instrument is not isolating the thing you hoped it was isolating.

A third approach is to model the intervening mechanism. If the treatment affects the outcome only through a set of measured intermediate variables, then we can use those variables to recover the causal effect. This is related to the front-door criterion. In practice this can be very hard, because the mechanism has to be exhaustive. If proppant affects production through stimulated rock volume, pressure changes, conductivity, pipe constraints, and a dozen other physical processes, then we need to measure the relevant parts of that chain well enough that there is no unmodeled path left over.

This is the part where it is tempting to say "well, let's just put everything in a regression and see what happens." That temptation is understandable. It is also how you get an answer that looks precise and is wrong.

Consider a simplified version of unconventional well production. Imagine that oil production is determined by geology, water saturation, and the size of the completion. Operators do not choose completion sizes randomly. In better parts of the basin, where porosity is higher and water saturation is lower, they tend to pump more sand because the expected return on that completion spend is better. In worse parts of the basin, they tend to pump less sand because there is less oil to contact in the first place.

If we create fake data from this process, we can know the true answer because we made the world ourselves. Suppose the true average treatment effect of adding proppant is about 1.4 barrels per foot for each additional 100 pounds of sand. The treatment effect is heterogeneous: in the best rock it might be closer to 2.5 barrels per foot, while in worse rock it might be closer to 0.7. That heterogeneity is important because the treated wells are not evenly distributed across rock quality. They are more likely to be in good rock.

Now fit the naive model. Predict oil per foot from proppant, porosity, and water saturation using a multivariate regression. The coefficient on proppant might come out around 2.4 barrels per foot, which is close to the effect in the best acreage and very far from the basin-wide average effect we were trying to estimate. The model has not discovered that proppant is twice as useful as we thought. It has confused part of the value of better rock with the value of a larger completion.

This matters because the business decision is different. If you overestimate the value of proppant, you will spend too much money on larger completions in places where the rock cannot reward that spend. If you underestimate the value of geology, you will attribute production gains to operational choices that were actually consequences of being in better acreage. The model can have good predictive performance and still give you the wrong answer to the decision question.

Stratification fixes the toy version of the problem. Separate the wells into acreage tiers and estimate the proppant effect inside each tier. The tier-one estimate is high, the tier-three estimate is low, and the tier-two estimate is in the middle. If we then take a weighted average using the prevalence of each tier in the basin, we recover the average treatment effect. We did not need a more impressive model. We needed the right comparison.

The same toy problem can also be attacked with a natural experiment. If an external sand shortage causes some wells to receive less proppant than operators otherwise would have chosen, and if that shortage is unrelated to geology and production capacity, then the shortage gives us variation in treatment that was not created by operator preference. Regressing production on that induced variation can recover a reasonable estimate of the average treatment effect.

The mechanistic approach is possible too, but the toy version reveals a different issue. Even if the causal structure is right, the statistical model can still be too simple. If the relationship between proppant, contacted hydrocarbons, and production is nonlinear, then a chain of linear regressions can underestimate the treatment effect for ordinary model-capacity reasons. Causal identification is not a substitute for modeling the phenomenon well. It only tells you which comparisons are meaningful.

This is one of the points that gets lost when causal inference is presented as a bag of methods. The method is not the hard part. There are packages in Python and R that implement many of these estimators, including generalized random forests, double machine learning, structural equation models, and graph-based causal workflows. The hard part is deciding what world you think produced the data. Which variables cause treatment assignment? Which variables affect the outcome? Which variables are colliders that should not be conditioned on? Which assumptions are plausible, and which assumptions are just wishes with Greek letters?

You cannot get that from the dataset alone. There are causal discovery methods that try to learn graph structure from data, but they still require assumptions, and they are not a replacement for understanding the system. The title of the webinar came from a Judea Pearl quote I like: if you are not smarter than your data, you should not be doing this in the first place. That sounds harsh, but I think it is mostly comforting. We are smarter than our data. We know things about how wells are drilled, how operators behave, how social networks transmit exercise habits, how alcohol consumption is bundled with other behaviors, and how business processes create selection effects. The data do not know those things unless we put them there.

There are also assumptions that are easy to forget because they are inconvenient. One important one is sometimes called stability under treatment values. The idea is that one unit's treatment assignment should not change another unit's treatment effect. This fails all the time. If I intervene on a group of people and get them to exercise more, their friends may start exercising too. If I drill wells closer together, the treatment applied to one well can affect the production of nearby wells. If a college degree becomes much more common while the number of jobs requiring one does not increase at the same rate, the value of the degree can change because other people also received the treatment.

This is especially relevant in unconventional development because wells are not independent marbles in a bag. They share rock, pressure regimes, infrastructure, parent-child relationships, depletion histories, and operational constraints. A model that treats them as independent observations can still be useful, but the assumptions behind a causal estimate need to be examined carefully. Downspacing, completion design, geology, and parent depletion are all tangled together because operators make coordinated development decisions in response to what they know about the basin.

This is also why causal inference is useful there. A normal predictive model may dramatically underestimate the impact of spacing by attributing the signal to geology. Better acreage gets tighter development, tighter development changes interference, and the observed production outcome is downstream of both. If you only ask the model to predict production, it can use whatever signal is convenient. If you ask how production would change under a different spacing design, you need something more disciplined.

The same idea shows up in redevelopment. Suppose an operator refracs an older well. We observe the production after the refrac, but the interesting quantity is incremental production: what did the refrac add above what the well would have produced if nobody had touched it? In that case the missing counterfactual is not just philosophical. It is the object we need for the economic calculation. We can forecast the no-intervention production path, compare it to the observed production path, and estimate the gain from the intervention. Better candidates tend to be wells where the original completion was small or stages were far apart, but the size of the gain also depends on where the well is in the basin.

The general pattern is the same across all of these examples. The decision-maker does not really care whether a feature is predictive. They care what will happen if they choose differently. Should we increase proppant loading? Should we drill tighter spacing? Should we refrac this well? Should we change the product experience? Should we recommend one medical treatment instead of another? These are causal questions. Predictive models can inform them, but they do not answer them automatically.

There is a tendency in data science to treat more flexible models as the solution to every estimation problem. If linear regression is biased, use a random forest. If the random forest is not good enough, use gradient boosting. If that is not good enough, use a neural network. This can work beautifully when the problem is prediction and the future looks like the past. It is much less reassuring when the question involves an intervention, because the intervention changes the data-generating process. The model is being asked about a world it did not observe.

That does not mean causal inference is magic. It is probably more honest to say that it makes the argument explicit. A causal estimate is only as good as the assumptions that identify it, the measurements that support those assumptions, and the model used to estimate the relevant quantities. Hidden confounders remain a problem. Measurement error remains a problem. Bad functional forms remain a problem. Interference remains a problem. The benefit is that causal inference gives us language and tools for saying which problem we think we have, rather than hiding it inside feature importance.

For a less technical introduction, Judea Pearl's The Book of Why is a good starting point. For practitioners who want the theory in more detail, Stephen Morgan and Christopher Winship's Counterfactuals and Causal Inference is a better fit. The software ecosystem is also much better than it used to be, with packages like DoWhy, EconML, and other toolkits making it possible to implement many of these estimators without writing everything from scratch.

But the software should be the last step, not the first one. Start by drawing the graph. Decide what the treatment is. Decide what outcome you care about. Decide which variables happened before the treatment and which happened after it. Decide where selection enters the system. Decide which assumptions you are willing to defend to someone who knows the domain well and is in a bad mood.

Only then should you estimate the effect. The dataset is not going to do that thinking for you.