# Anyone interested in practicing hands-on causal inference?

**URL:** <https://community.intuitivebayes.com/t/anyone-interested-in-practicing-hands-on-causal-inference/641>\
**Category:** Causal Inference Book Club\
**Tags:** exercises\
**Created:** [December 18, 2022, 3:56pm UTC](https://community.intuitivebayes.com/t/anyone-interested-in-practicing-hands-on-causal-inference/641 "2022-12-18T15:56:15Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![rahulB](https://yyz1.discourse-cdn.com/flex007/user_avatar/community.intuitivebayes.com/rahulb/32/140_2.png) [@rahulB](https://community.intuitivebayes.com/u/rahulB)\
**Post date:** [December 18, 2022, 3:56pm UTC](https://community.intuitivebayes.com/t/anyone-interested-in-practicing-hands-on-causal-inference/641/1 "2022-12-18T15:56:15Z")

</div>

Hi All,

I was thinking of getting some hands-on experience in causal inference by applying some of the learnings on real-world datasets. There are some good open and anonymized datasets from companies available on scikit-uplift package website [here](https://www.uplift-modeling.com/en/latest/tutorials.html#basic). The [X5\_RetailHero](https://www.uplift-modeling.com/en/latest/api/datasets/fetch_x5.html#x5) in particular looks really interesting. Moreover, this dataset was a part of a [competition](https://ods.ai/competitions/x5-retailhero-uplift-modeling/data) held some 2 years ago.  
Would anyone be interested in trying this out in the next few weeks before beginning with the new topic?

---

<div class="post-metadata">

**Author:** ![RavinKumar](https://yyz1.discourse-cdn.com/flex007/user_avatar/community.intuitivebayes.com/ravinkumar/32/13_2.png) [@RavinKumar](https://community.intuitivebayes.com/u/RavinKumar)\
**Post date:** [December 18, 2022, 4:17pm UTC](https://community.intuitivebayes.com/t/anyone-interested-in-practicing-hands-on-causal-inference/641/2 "2022-12-18T16:17:37Z")

</div>

This is a great idea. More hands on practice will reinforce the concepts we’ve learned. I’ve been kicking around some ideas as well.

What format were you thinking?

---

<div class="post-metadata">

**Author:** ![ChadDelany](https://yyz1.discourse-cdn.com/flex007/user_avatar/community.intuitivebayes.com/chaddelany/32/78_2.png) [@ChadDelany](https://community.intuitivebayes.com/u/ChadDelany)\
**Post date:** [December 18, 2022, 5:35pm UTC](https://community.intuitivebayes.com/t/anyone-interested-in-practicing-hands-on-causal-inference/641/3 "2022-12-18T17:35:50Z")

</div>

I’m very interested. I’ve been looking for different datasets to do some diff-in-diff analysis on.

---

<div class="post-metadata">

**Author:** ![rahulB](https://yyz1.discourse-cdn.com/flex007/user_avatar/community.intuitivebayes.com/rahulb/32/140_2.png) [@rahulB](https://community.intuitivebayes.com/u/rahulB)\
**Post date:** [December 18, 2022, 8:04pm UTC](https://community.intuitivebayes.com/t/anyone-interested-in-practicing-hands-on-causal-inference/641/4 "2022-12-18T20:04:34Z")

</div>

These datasets are obtained from randomized control experiments and as such one can simply calculate `ATE = Y_1 - Y_0`. But, we can treat these as observational studies and apply the causal inference methods such as matching, sub-classification, DiD etc. to see how close we can come to the real ATE. We can also compute heterogeneous treatment effects for individual users. I was thinking of

1. Manually computing ATE, ATT using matching methods or IPW etc.
2. Using the libraries `dowhy` and `econml` from Microsoft to compare the values calculated manually.
3. Compare methods to see which one does the best.

I do not have a strong opinion on the format. I was thinking of starting a public git repo and put my code in a folder under my name. Others can refer it, or, fork the repo and create a folder under their name and subsequently create a pull request to main. This way all code is in one place and everyone can refer. But, **please feel free to suggest other methods that you think would be better.**  
It would be interesting to talk about the methods others apply (there is always some degree of subjectivity to causal analysis) and then talk about the results in the next 2-3 weeks time.

Let me know how this sounds and any suggestions or comments are welcome.
