We have undergone some system upgrade and now we use the latest pandas version 3.0.5. Crema used to work with the older versions, but now it fails with the newer version of panda, but with the old data.
This is the suggestions from Claude:
crema.parsers.txt._convert_target_col() detects a target/decoy column that holds strings (e.g. Tide's "target"/"decoy" values) by checking:
elif data.dtype == "object":
Starting with pandas 3.0, columns of Python strings read via pd.read_csv no longer default to the legacy object dtype — pandas now infers a dedicated string dtype (reported as str/string[pyarrow]) by default. Because that dtype isn't "object", this branch is skipped and execution falls through to:
return data > 0
which raises, since the column still holds strings:
TypeError: Invalid comparison between dtype=str and int
Traceback (via crema.read_tide → read_txt → _convert_target_col):
File ".../crema/parsers/txt.py", line 108, in read_txt
data[target_column] = _convert_target_col(data[target_column])
File ".../crema/parsers/txt.py", line 217, in _convert_target_col
return data > 0
TypeError: Invalid comparison between dtype=str and int
Environment: Python 3.14, pandas 3.x (the new default-string-dtype behavior; introduced under pandas.options.future.infer_string in 2.x, made the default in 3.0).
Suggested fix:
Check for string-like dtype in general, not just object, e.g.:
elif data.dtype == "object" or pd.api.types.is_string_dtype(data):
targets = {
"target": True,
"t": True,
"decoy": False,
"d": False,
"f": False,
}
return data.astype(object).map(targets)
pd.api.types.is_string_dtype covers both the legacy object-backed strings and the new pandas string dtype, so this works across pandas versions.
Want me to also open this as an actual issue on the repo, or just hand you the text to post yourself?
We have undergone some system upgrade and now we use the latest pandas version 3.0.5. Crema used to work with the older versions, but now it fails with the newer version of panda, but with the old data.
This is the suggestions from Claude:
crema.parsers.txt._convert_target_col() detects a target/decoy column that holds strings (e.g. Tide's "target"/"decoy" values) by checking:
elif data.dtype == "object":
Starting with pandas 3.0, columns of Python strings read via pd.read_csv no longer default to the legacy object dtype — pandas now infers a dedicated string dtype (reported as str/string[pyarrow]) by default. Because that dtype isn't "object", this branch is skipped and execution falls through to:
return data > 0
which raises, since the column still holds strings:
TypeError: Invalid comparison between dtype=str and int
Traceback (via crema.read_tide → read_txt → _convert_target_col):
File ".../crema/parsers/txt.py", line 108, in read_txt
data[target_column] = _convert_target_col(data[target_column])
File ".../crema/parsers/txt.py", line 217, in _convert_target_col
return data > 0
TypeError: Invalid comparison between dtype=str and int
Environment: Python 3.14, pandas 3.x (the new default-string-dtype behavior; introduced under pandas.options.future.infer_string in 2.x, made the default in 3.0).
Suggested fix:
Check for string-like dtype in general, not just object, e.g.:
elif data.dtype == "object" or pd.api.types.is_string_dtype(data):
targets = {
"target": True,
"t": True,
"decoy": False,
"d": False,
"f": False,
}
return data.astype(object).map(targets)
pd.api.types.is_string_dtype covers both the legacy object-backed strings and the new pandas string dtype, so this works across pandas versions.
Want me to also open this as an actual issue on the repo, or just hand you the text to post yourself?