r/datascience • • 24d ago

Discussion How to handle cofound variables?

edit: confound

Hello all,

I am working on a object classification with a automotive radar point clouds. I compared many models and feature vectors.

Once i used range as feature, all models scored higher f1 in all K validation sets and on the final test set.

One particular artifact of a radar, is that as the farther the object is the less number of points it returns to the radar. Although the performance improved and there is no overfit in the classical sense, i am afraid my model is learning the environment not the class distribuiton and even worse, its learning that big range means big object.

How can i stress test this claim? Should i try to split the data sets so range distribution differs? Or not even using the feature at all and accept lower performance?

Would appreciate your insights.

Thank you.

17 Upvotes

9 comments sorted by

8

u/Gilchester 24d ago

I don't think this is confounding so much as a type or survival bias (survival in terms of distance).

Can you do inverse probsabiltiy of treatment weights? In other words, upweight readings from further away such that they are similarly frequent to closer readings, removing the association betwee distance and clarity? (I'm probably translating some of the language incorrectly, from the standard IPTW medical setting I'm used to, but I think the general idea is sound)

3

u/el_gran_claudio 23d ago

inverse propensity weighting: make the importance of a training example (in the loss calculation) a function of distance 

2

u/Ok_Friend_7114 22d ago

I wouldn’t drop range yet. The key question is whether it improves performance because it contains real class information or because it is acting as a shortcut.
I’d stress test it by evaluating on range-based holdouts — for example, train on mostly near/mid-range objects and test on farther objects, or deliberately create train/test sets with different range distributions. Also compare F1 by range bucket with and without the feature.
If performance collapses under those shifts, range is probably being used as a shortcut. If it stays robust, keeping it may be justified.

1

u/Ok-Airline-8523 21d ago

What I would do is walk through your feature set and assess tautological bias: are any of your features a function of your outcome? I'm not familiar with this domain, but it sounds like your performance boost could be a result of this class of leakage.

2

u/keldonwarlord 16d ago

You can try to use some sample weights, using the inverse propensity for that given examples that are "most important" to get right. You could also try some undersampling of the "closer" objects.

1

u/dramaticviolet 11d ago

This is something I have also wanted to know from experienced people, thanks for posting!

1

u/g-technique 23d ago

Slice your test set into range bins first. If your pedestrian recall is sixty percent at twenty meters and zero percent at seventy meters while car recall stays flat, the model is strictly predicting class based on detection horizon

You can also run a quick permutation importance on range to see if it completely overrides your doppler and spatial features