A reproducible benchmark on the public 7M-row Criteo uplift dataset: S-, T-, X-, DR-Learner and Causal Forest against a naive response-ranking baseline, all judged with AUUC-based validation.
S/T/X/DR-Learner, Causal Forest, and a naive baseline, head to head.
Run at real scale on the public Criteo dataset.
Judged by uplift-appropriate metrics, not raw accuracy.
I ship it tested, documented, and ready to run in your stack.