[D] Pregnancy in a multi gender model

Gadod_mais@alien.top · 1 year ago

[D] Pregnancy in a multi gender model

tw_f@alien.top · 1 year ago

I’d just set it to False when sex=“M” as men can’t get pregnant and if you have a True value anywhere in that column it’s probably an annotation error.

This is the most pragmatic solution I guess.

314per@alien.top · 1 year ago

Keep in mind that it is possible but rare for men to be pregnant (eg trans men, some kind of intersex people).

https://www.healthline.com/health/transgender/can-men-get-pregnant#if-you-have-a-uterus-and-ovaries

Pyrite_Pro@alien.top · 1 year ago

Trans men, you mean… women?

pan_berbelek@alien.top · 1 year ago

And of course you get downvoted… the virus devours peoples minds.

Pyrite_Pro@alien.top · 1 year ago

Didn’t expect anything else…

liquidInkRocks@alien.top · 1 year ago

It’s been nice having you on Reddit, but prepare to be banned. They don’t tolerate any variance from the Progressive mindset.

UnusualClimberBear@alien.top · 1 year ago

Usually adding a L1/L2 regularization is an efficient way to deal with correlated variables

VoidRippah@alien.top · 1 year ago

The purpose of the model is to predict whether people infected with COVID will die or not.

There is no need for all that effort, I can tell you with 100% accuracy that they will day for certain. Just like those who where not infected. /s

grudev@alien.top · 1 year ago

TIL: ML redditors don’t understand sarcasm or the downsides of “accuracy” in model evaluation.

mathbbR@alien.top · 1 year ago

As others have said, pregnant men, while uncommon, might appear in your dataset. However, for preliminary results, it could make sense to exclude the possibility. Depending on the prevalence and impact of pregnancy in your data, you may be free to drop the column entirely. Some exploratory data analysis should help. Pregnancy should be relatively frequent (>5% of rows, at the very least). Pregnant women should also have noticably different target statistics than non-pregnant women. If pregnancy is rare or doesn’t seem to have much of an impact, feel free to drop it altogether for now.

Alternatively, I would default to basic indicator columns for both sex and pregnancy. Just be sure your model doesn’t require feature independence, as some do.

deeceeo@alien.top · 1 year ago

It likely won’t matter. Most models (e.g. I’m guessing something like xgboost) can deal robustly with these types of correlations.

If you like, you can combine the two into a single variable and may get slightly improved performance (0 for male, 1 for female and 2 for pregnant female) assuming the dataset can fit the rule (e.g. trans men). This way, a tree-based model could draw a boundary between 0 and 1 based on gender or 1 and 2 based on pregnancy.