Skip to content

P1 regression EDA crashes on skew/kurtosis if nan present in outcome_label mixed-type outcomes (e.g. NaN etc.) #9

Description

@raptor419

P1 regression EDA crashes on skew/kurtosis if nan present in outcome_label mixed-type outcomes

Problem:

Phase 1 could crash during initial EDA for regression/continuous outcomes when the outcome column contained mixed types, missing values, infinities, or non-numeric strings.

Image

TypeError: ufunc 'isnan' not supported for the input types, and the inputs could not be safely coerced to any supported types according to the casting rule "safe"

Cause:

scipy.stats.skew() and scipy.stats.kurtosis() were being called directly on the outcome column. When the column had object/mixed dtype, SciPy’s NaN handling could fail before the values were safely converted to numeric form.

Fix:

The regression EDA skew/kurtosis calculation now cleans the outcome values before calling SciPy. It coerces the outcome column to numeric, replaces infinities with missing values, drops missing/non-numeric values, and computes skewness/kurtosis only when valid numeric values remain. If no numeric values remain, STREAMLINE logs that skewness/kurtosis are unavailable.

Result:

P1 now continues gracefully for regression datasets with mixed-type outcome columns instead of crashing during initial EDA.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions