Summary
PepQuery 2.0.2, run in novel peptide validation mode (-s 1) with a
downloaded reference database (-db gencode:human), aborts while building the
protein-inference index:
java.lang.IllegalArgumentException: Protein sequence of protein
'ENSP00000402491.1 ENSP00000402491.1|ENST00000457194.6|ENSG00000211454.16|OTTHUMG00000002520.6|OTTHUMT00000127796.3|AKR7L-202|AKR7L'
contains invalid character: '*'.
at com.compomics.util.experiment.identification.protein_inference.fm_index.FMIndex.addDataToIndex(FMIndex.java:1454)
at com.compomics.util.experiment.identification.protein_inference.fm_index.FMIndex.init(FMIndex.java:1204)
at com.compomics.util.experiment.identification.protein_inference.fm_index.FMIndex.<init>(FMIndex.java:590)
at main.java.pg.PepMapping.init(PepMapping.java:81)
at main.java.pg.PepMapping.loadDB(PepMapping.java:47)
at main.java.pg.InputProcessor.searchRefDB(InputProcessor.java:419)
at main.java.pg.PeptideSearchMT.search(PeptideSearchMT.java:337)
at main.java.pg.PeptideSearchMT.search_multiple_datasets(PeptideSearchMT.java:738)
at main.java.pg.PeptideSearchMT.main(PeptideSearchMT.java:183)
at main.java.pg.IMain.main(IMain.java:30)
The reference database here is the one PepQuery downloads itself from the
gencode:human shorthand, not a user-supplied file.
What distinguishes a failing run from a working one
Same -db gencode:human, same MS dataset, same peptide input — only the
search type differs:
-s |
mode |
result |
1 |
novel peptide validation |
aborts as above |
2 |
known peptide validation |
completes normally |
That fits the stack: -s 1 reaches InputProcessor.searchRefDB →
PepMapping.loadDB, which builds the FMIndex over the reference; -s 2
does not.
Supplying a reference FASTA directly (-db /path/to/some.fasta) also works,
which is consistent with most user-supplied FASTAs not containing *.
Why the reference is not at fault
GENCODE translation FASTAs legitimately use * to mark stop codons — it is
part of the format, not corruption. So the downloaded file is valid and the
indexer is rejecting a character the source format permits. AKR7L-202 is
simply the first sequence in the file where one appears.
Possible directions:
- Strip or substitute
* (and any other non-standard residue) when loading
sequences into the FM index.
- Skip and count offending sequences rather than aborting, with a warning.
Either would make the built-in gencode:human shorthand usable with -s 1,
which is currently the combination that fails.
Secondary: the process exits 0
The exception is fatal and unhandled, but the process still returns exit
code 0. Anything running PepQuery non-interactively has to scrape stderr to
notice, rather than trusting the exit status. Returning non-zero on an
unhandled exception would be worth having independently of the fix above.
Environment
- PepQuery 2.0.2, the bioconda package built from the
pepquery.org 2.0.2
tarball (the latest bioconda offers)
- Invoked through the Galaxy
pepquery2 wrapper, which passes these arguments
through unchanged; the discriminating arguments are -s 1|2,
-db gencode:human, -t peptide. I have not run the equivalent bare
command line myself, so I have quoted the arguments rather than reconstruct
a full invocation.
- Reproducible, not intermittent: the two wrapper tests using
-s 1 with this
database failed on every one of the last six nightly runs (6/6 and 4/4
executions that completed), while the -s 2 test using the same database
passed every time.
Galaxy-side tracking issue, for cross-reference:
galaxyproteomics/tools-galaxyp#840
Summary compiled by Claude (Claude Code), working with @afgane.
Summary
PepQuery 2.0.2, run in novel peptide validation mode (
-s 1) with adownloaded reference database (
-db gencode:human), aborts while building theprotein-inference index:
The reference database here is the one PepQuery downloads itself from the
gencode:humanshorthand, not a user-supplied file.What distinguishes a failing run from a working one
Same
-db gencode:human, same MS dataset, same peptide input — only thesearch type differs:
-s12That fits the stack:
-s 1reachesInputProcessor.searchRefDB→PepMapping.loadDB, which builds theFMIndexover the reference;-s 2does not.
Supplying a reference FASTA directly (
-db /path/to/some.fasta) also works,which is consistent with most user-supplied FASTAs not containing
*.Why the reference is not at fault
GENCODE translation FASTAs legitimately use
*to mark stop codons — it ispart of the format, not corruption. So the downloaded file is valid and the
indexer is rejecting a character the source format permits.
AKR7L-202issimply the first sequence in the file where one appears.
Possible directions:
*(and any other non-standard residue) when loadingsequences into the FM index.
Either would make the built-in
gencode:humanshorthand usable with-s 1,which is currently the combination that fails.
Secondary: the process exits 0
The exception is fatal and unhandled, but the process still returns exit
code 0. Anything running PepQuery non-interactively has to scrape stderr to
notice, rather than trusting the exit status. Returning non-zero on an
unhandled exception would be worth having independently of the fix above.
Environment
pepquery.org2.0.2tarball (the latest bioconda offers)
pepquery2wrapper, which passes these argumentsthrough unchanged; the discriminating arguments are
-s 1|2,-db gencode:human,-t peptide. I have not run the equivalent barecommand line myself, so I have quoted the arguments rather than reconstruct
a full invocation.
-s 1with thisdatabase failed on every one of the last six nightly runs (6/6 and 4/4
executions that completed), while the
-s 2test using the same databasepassed every time.
Galaxy-side tracking issue, for cross-reference:
galaxyproteomics/tools-galaxyp#840
Summary compiled by Claude (Claude Code), working with @afgane.