Assignment 5 - Ifs, Onlys, and Buts
This assignment gets you conceptual practice with the last layers of the causal hierarchy, as well as some linear SCMs, practical counterfactuals, and some programmatic practice too!
This is a Group Assignment! Feel free to work in groups up to the group-size limit listed in the syllabus.
Solution Skeleton
Start with the solution skeleton in-hand! In the following project, I've given you only the desired directory structure required for the problems that follow -- the rest is up to you!
Inside you'll find the directories:
docin which you should place your finalreport.pdfcontaining your answers to every question requesting a written response, computation, etc. I suggest you compose this document using Markdown, Word with its equations editor, or LaTeX (preferably). Enumerate your answers corresponding to each problem, clearly indicating what answer corresponds to what problem!srcalthough not required, if you had any scripts used to save yourself some by-hand computation, you can use these instead of performing calculations by hand (place all scripts here).
Specifications
1337 Hiring
Let's start with a familiar problem because I can't be bothered to make another!
Counterfactuals and Fully-Specified SCMs: We're examining the hiring practices of a particular organization that has a reputation for being elitist and rampantly afflicted by croneyism (*Andrew puts down his word-of-the-day calendar*). From the company's hiring data and reasonable assumptions about the causal structure about the system, you're able to specify some variables of interest:
\(S \in \{0,1\}\), the socio-economic status of the job candidate with \(0\) being low-SES, \(1\) being high.
\(E \in \{0,1\}\), the education level of the candidate with \(0\) being low-status (e.g., ill-reputed undergrad institution or low GPA) and \(1\) being high.
\(H \in \{0,1\}\), whether or not that candidate was hired with \(0\) being not.
\(U_E \in \{0,1\}\)*, exogenous errors lumping in all the reasons someone in each SES may or may not go to college, like opting instead to pursue a career as a Twitch streamer.
\(U_H \in \{0,1\}\)*, exogenous errors lumping in all the reasons someone may or may not be hired, like being a college roommate of the CEO, despite being dumb as a rock.
* Note: often, exogenous variables are seen as ways to model "normalcy" such that the system operates according to known influences *except for* when the exogenous variables represent an unmodeled exception (see structural equations demonstrating this below).
The model for this system (including structural equations for endogenous variables):
Your task: Compute the probability of necessity of attending an elite college to get hired, \(P(H_{E=0}=0|H=1,E=1)\), through the following steps:
Draw the twin network model associated with this query.
Show which CPTs, if any relevant for the query, would be updated by the abduction step of structural counterfactuals, and to what values (show your steps of enumeration inference).
You may then solve for the query by combining the action + prediction steps, showing your work.
Up and Down the Ladder
Causal Hierarchy Dynamism: we've already seen some tools for dancing up and down the causal ladder, taking queries that we might have at one tier, and being able to answer them with data from a lower. Here are a couple of extra ones that are handy.
z-Specific Effects are causal effects measured in the context of a specific subpopulation like "The effect of Drug X=1 on recovery Y=1 from migraines amongst older users Z=1, 65 and older:" \(P(Y=1|do(X=1), Z=1)\)
However, because we are now including observed evidence \(Z\) as part of the query, we need to make sure that this addition of evidence doesn't create any spurious pathways that we must block to answer the interventional \(do(X)\) component!
This observation gives us a modified, somewhat more general, version of the back-door criterion.
The z-specific effect of X on Y is identified whenever we can find a set of variables \(S\) such that the set \(S \cup Z\) (where \(\cup\) is the set union operator) satisfies the back-door criterion. In the case where such a set can be found, the adjustment formula for some pre-treatment contexts \(Z\) (i.e., if \(Z\) are non-descendants of \(X\)) is: $$P(Y=y|do(X=x), Z=z) = \sum_s P(Y=y|X=x, Z=z, S=s) P(S=s | Z=z)$$
z-specific effects are found somewhat at the union of tiers 1 and 2 but not quite at tier 3 where evidence is allowed to clash with intervention. That said, the above adjustment again gives us the ability to express the tier-2 component in the query \(P(Y=1|do(X=1), Z=1)\) in terms of strictly tier-1 expressions.
In that same "stepping-down the ladder" direction, there is also an interpretation of the back-door theorem that allows us to step from tier-3 to tier-2.
Counterfactual Backdoor: If a set Z of variables satisfies the backdoor criterion relative to (X, Y), then, for all x, the counterfactual \(Y_x\) is conditionally independent of (tier-1) X given (tier-1) Z: $$P(Y_{X=x} | X=x', Z=z) = P(Y_{X=x} | Z=z) = P(Y | do(X=x), Z=z)$$
Intuitively, the above says that, if we block a path between the observational and experimental models in our twin network, then the counterfactual query is really just a causal query. To wit, remember our twin network for the infamous firing squad:
Suppose, in the above, we wished to estimate \(P(D_{S_1 = 0}|S_1 = 1, C=1)\), then we see that the observed information of \(S_1=1\) is blocked from "crossing over the bridge" through \(G\) by knowing \(C=1\). This means that, by the counterfactual backdoor: $$P(D_{S_1 = 0}|S_1 = 1, C=1) = P(D|do(S_1=0),C=1)$$ ...thus turning a tier-3 query into a tier-2. Neat!
SO, to get some practice with the above, let's work with the following model:
Your tasks: for each of the following queries at tiers 2 and 3 of the hierarchy, use the above to give the adjustment formula that would be used to compute each in tier-1 terms alone.
Specifically: (1) state which adjustment formula you're using, and, if applicable, the choice of set \(S\) and / or \(Z\), then (2) derive the adjustment formula to express each query below in tier-1 expressions only.
\(P(W|do(Z), Y)\)
\(P(Y|do(X), W)\)
\(P(W_{Z=z}|Z=z', Y=y, X=x)\)
ETT Tu?
Effect of Treatment on the Treated: You're joining a new web advertising agency as a replacement for Dave. Everyone loves you by default because everyone hated Dave. He was the type to add gratuitous memes to his meeting PowerPoints, and wanted to constantly discuss the latest episode of Homeland. Worse yet, he was the only one nursing a legacy advertising agent (written entirely in C) that no one else knows how to operate, what user features it employs, nor if it is particularly effective at translating ads to clickthroughs. Your job is to recreate a more interpretable ad agent, but to first make sure that the company isn't hemorrhaging money from bad advertising due to Dave's lame agent.
Suppose Dave's agent currently works by mapping some arcane user demographics to one of two ad portfolios \(X \in \{0,1\}\) that will display targeted ads to users, after which it is recorded whether or not the user clicked on the ad \(Y \in \{0,1\}\). Sadly, the system is affected by confounding, since you do not know how or what demographics are considered by the agent, nor how they affect the clickthrough rate. However, as a quick means of determining how good Dave's system is, you seed a small propensity for the agent to randomly choose an ad portfolio to show to the user, and collect the following data:
Your tasks:
Determine (by providing rationale and evidence for) whether or not Dave's agent is operating optimally. As part of your argument, determine the percentage of clickthroughs \((Y=1)\) that the agent is either missing or achieving above a baseline random agent (i.e., an agent that randomly chooses an ad for each viewer). You must derive this using the tools of causal inference and using the correct notation that spans the right tiers of the causal hierarchy.
Is it possible, without changing Dave's agent, to create a second agent that performs better than Dave's even without knowing what the unobserved confounders are? Describe how you could deploy this second agent if so.
Hint: Think about making an agent that takes Dave's agent's decision as a parameter, i.e. $$\pi_{you}(\pi_{Dave}(s), s) = \text{action_better_than_}\pi_{Dave}(s)$$
Intent Specificity
Empirical Counterfactuals: You are analyzing the results of a Heterogeneous-Intent Randomized Clinical Trial (HI-RCT) studying recovery \(Y \in \{0,1\}\) rates of some drug treatments \(X \in \{0,1,2\}\) in a diagnostic setting wherein physicians are subject to unobserved confounding factors (for simplicity, assumed to be all exchangeable in their intent-formation, meaning they're all sensitive to any UCs in the same way). However, due to sampling constraints, you are only able to collect the following data with some missing entries indicated by [?]:
Your tasks: (SHOWING YOUR WORK for each)
Using the information in the tables above, solve for \(P(Y=1 | X=1)\) (Hint: remember some of the axioms of counterfactual notation).
Using the information in the tables above, solve for \(P(Y_{X=2}=1 | X=0)\).
If you were building a recommender system that served as a "driver assist" for physician treatments, such that each physician entered their intended treatments \(I \in \{0,1,2\}\), what drug would your system recommend for each possibly intended drug?
Hint: think about the Dave situation in the problem above, except instead of Dave's policy, you use the physician's intended treatment.
Attribution and Necessity
Attribution and Necessity: The legal criterion of "but for" is summarized by the counterfactual probability of necessity, according to which judgment in favor of a plaintiff should be made if and only if it is "more probable than not" that the damage would not have occured "but for" the defendant's actions. Expressing this probability for some action \(X\) and outcome \(Y\) as: $$PN(x,y) = P(Y_{X=x'}=y' | X=x, Y=y)$$
It turns out that, even in absence of an SCM, but with both observational and experimental data, the PN can be estimated with only mild assumptions (known as monotonicity if Y is monotonic relative to X such that \(Y_{X=1}(u) \ge Y_{X=0}(u)~\forall~u\)). This is done through 2 tools:
The Excess Risk Ration (ERR) is a variant of Risk Difference typically used in court cases in the absence of experimental data; it is also known as the Attributable Risk Fraction among the Exposed, and is written (for outcome \(y\) and treatment/action \(x\)): $$ERR(x, y) = \frac{P(y|x) - P(y|x')}{P(y|x)}$$
The Confounding Factor (CF) is a correction needed to account for confounding bias in the ERR, i.e., whenever the statistical test for confounding yields \(P(y|do(x')) \ne P(y|x')\), and is defined as: $$CF(x, y) = \frac{P(y|x') - P(y|do(x'))}{P(x, y)}$$
Suppose there is a case brought against a car manufacturer claiming that its car's faulty design led to a man's death in a car crash. The ERR tells us how much more likely people are to die in crashes when driving one of the manufacturer's cars, but if it turns out that people who buy their cars are more likely to drive fast (leading to deadlier crashes) than the general population, the CF corrects for it.
Interestingly, the PN can be rephrased as: $$PN(x, y) = ERR(x, y) + CF(x, y)$$
With that in mind, let's put our lawyer hats on!
A lawsuit is filed against a Big Pharma company that produces Drug X, alleging that the drug is likely to have caused the death of Mr. Cy A. Nide, who took it to relieve back pains. The manufacturer claims that experimental data on patients with back pains show conclusively that Drug X only has minor effects on mortality, but the plaintiff argues that this population data is only good for averages, not for patients like Mr. Nide who did not participate. To support her claim, the plaintiff displays results of a nonexperimental survey of patients who suffered back pain and took Drug X, the results of which are below:
Your task: As the court, determine whether it is "more probable than not" that the drug was responsible for Mr. Nide's death.
VHD Adieu
Did you really think you'd escape the semester without one last visit from the Vaping-Heart-Disease (VHD) network? This time, it's getting... linear.
In the above:
Ignoring how we found them (consult our causality textbook), the "path coefficients" detailing how much of a causal effect each cause has per unit-increase on its effect. These are what supply the coefficients in each linear structural equation, e.g., \(V \leftarrow f_V(S, U_V) = 2*S + U_V\), where 2 is the path coefficient from \(S \rightarrow V\). In words, this means that, on average, \(V\) will increase by 2 every time \(S\) increases by 1.
Apropos, suppose that the units of each variable measured here are in standardized Z-scores, i.e., a Z-score of 0 means that the variable has its average value. E.g., if we say \(V = 0\), this means that someone vapes an average amount, but \(V = -1\) means that they vape 1 standard deviation less than the average individual.
As a result, and typical when modeling Z-scores, we'll assume that the expected value of all exogenous vars, \(U_X\), are 0: \(E[U_X] = 0~\forall~X\).
Your Task: using the LSCM and assumptions above, find (showing your steps): \(E[H | do(V=2), S=-2]\)
Submission
You will be submitting your assignments through GitHub Classroom!
What
Complete all requested sections of your report.pdf (making sure to clearly label problem numbers and answers) and place in your doc
directory. Complete all optional Python scripts and place in your src directory.
How
To clone this assignment (if you need a refresher), consult the guide here:
To submit this assignment:
Simply push your final, submission copy to the GitHub Classroom repository associated with your account.
Place your name at the top of *all* submitted files AND in the accompanying
readmefile.