Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
72 changes: 72 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
name: CI

on:
push:
pull_request:

permissions:
contents: read

jobs:
test:
name: Install & smoke-test (Python ${{ matrix.python-version }})
runs-on: ubuntu-latest

strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11"]

env:
# Stable cache key – bump the suffix if the architecture changes.
CKPT_CACHE_KEY: cascadia-dummy-ckpt-arch-v1

steps:
- uses: actions/checkout@v4

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Disable checkout credential persistence.

This workflow executes pull-request-controlled code via package installation, fixture generation, and pytest. actions/checkout@v4 otherwise stores GITHUB_TOKEN in .git/config, allowing malicious setup or test code to read and exfiltrate it. contents: read limits the token’s permissions but does not prevent credential exposure.

🔒 Proposed fix
       - uses: actions/checkout@v4
+        with:
+          persist-credentials: false
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
- uses: actions/checkout@v4
- uses: actions/checkout@v4
with:
persist-credentials: false
🧰 Tools
🪛 zizmor (1.26.1)

[warning] 25-25: credential persistence through GitHub Actions artifacts (artipacked): does not set persist-credentials: false

(artipacked)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/ci.yml at line 25, Update the actions/checkout@v4 step to
disable credential persistence by setting its persist-credentials option to
false, while leaving the existing checkout behavior unchanged.

Source: Linters/SAST tools


- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
cache: pip

# Resolve the checkpoint path with the real home dir so Python never
# receives a literal ~ that Path.mkdir() would treat as a directory name.
- name: Set checkpoint path
run: echo "CKPT_PATH=$HOME/.cache/cascadia/dummy.ckpt" >> "$GITHUB_ENV"

# Install PyTorch CPU-only first so pip does not pull in the multi-GB
# CUDA wheel when it later resolves cascadia's torch dependency.
- name: Install PyTorch (CPU-only)
run: |
pip install --upgrade pip
pip install torch --index-url https://download.pytorch.org/whl/cpu

- name: Install cascadia and its dependencies
run: pip install ".[test]"

# Quick import smoke-test before running the full suite.
- name: Verify import
run: python -c "import cascadia; print('cascadia imported OK')"

# Restore a previously generated dummy checkpoint. The key is stable
# (tied only to the model architecture version) so it is shared across
# branches and only regenerated when the architecture changes.
- name: Restore dummy checkpoint cache
id: ckpt-cache
uses: actions/cache@v4
with:
path: ${{ env.CKPT_PATH }}
key: ${{ env.CKPT_CACHE_KEY }}-py${{ matrix.python-version }}

- name: Generate dummy checkpoint (cache miss only)
if: steps.ckpt-cache.outputs.cache-hit != 'true'
run: python tests/create_dummy_ckpt.py "$CKPT_PATH"

# Run the full test suite. The augmentation tests exercise the mzML
# parsing and ASF generation path; the e2e test drives the complete
# ``cascadia sequence`` pipeline using the dummy checkpoint.
- name: Run tests
env:
CASCADIA_DUMMY_CKPT: ${{ env.CKPT_PATH }}
run: pytest tests/ -v --tb=short
1 change: 1 addition & 0 deletions cascadia/cascadia.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@
import argparse
from lightning.pytorch import loggers as pl_loggers
from .depthcharge.utils import *
from .utils import write_results
from .model import AugmentedSpec2Pep
from .augment import *
from datetime import datetime
Expand Down
7 changes: 5 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,9 @@ classifiers = [
requires-python = ">=3.8"
dependencies = [
"setuptools<70.0.0",
"lightning>=2.0,<2.1",
"pytorch-lightning>=1.9,<2.0",
"lightning>=2.0,<3.0",
"pyteomics>=4.6",
"psims",
"torch>=2.2.0",
"numpy<2.0",
"numba>=0.48.0",
Expand All @@ -34,6 +34,9 @@ dependencies = [
"tensorboard",
]

[project.optional-dependencies]
test = ["pytest>=7.0"]

[project.scripts]
cascadia = "cascadia.cascadia:main"

Expand Down
Empty file added tests/__init__.py
Empty file.
63 changes: 63 additions & 0 deletions tests/create_dummy_ckpt.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
"""Create a random-weight Cascadia checkpoint for CI end-to-end testing.

The checkpoint uses the *production* architecture (d_model=512, n_layers=9)
so that ``cascadia sequence`` can load it without modification. Weights are
randomly initialised with a fixed seed so the file is reproducible.

We write the Lightning checkpoint format directly with ``torch.save`` rather
than going through ``Trainer.save_checkpoint``, which requires the trainer to
have completed the internal fit-setup chain (unavailable without a training
run). ``LightningModule.load_from_checkpoint`` only requires ``state_dict``
in the checkpoint dict when all hyperparameters are supplied as kwargs.

Usage (called by the CI workflow):
python tests/create_dummy_ckpt.py <output_path>
"""
import sys
from pathlib import Path

import torch
import pytorch_lightning as pl


def create_dummy_ckpt(output_path: str | Path) -> Path:
output_path = Path(output_path).expanduser()
output_path.parent.mkdir(parents=True, exist_ok=True)

torch.manual_seed(0)

from cascadia.depthcharge.tokenizers import PeptideTokenizer
from cascadia.model import AugmentedSpec2Pep

tokenizer = PeptideTokenizer.from_massivekb(
reverse=False, replace_isoleucine_with_leucine=True
)

model = AugmentedSpec2Pep(
d_model=512,
n_layers=9,
n_head=8,
dim_feedforward=1024,
dropout=0,
rt_width=2,
tokenizer=tokenizer,
max_charge=10,
)

checkpoint = {
"epoch": 0,
"global_step": 0,
"pytorch-lightning_version": pl.__version__,
"state_dict": model.state_dict(),
"hyper_parameters": {},
}
torch.save(checkpoint, str(output_path))

size_mb = output_path.stat().st_size / 1e6
print(f"Saved dummy checkpoint to {output_path} ({size_mb:.1f} MB)")
return output_path


if __name__ == "__main__":
out = Path(sys.argv[1]) if len(sys.argv) > 1 else Path("dummy.ckpt")
create_dummy_ckpt(out)
153 changes: 153 additions & 0 deletions tests/create_mini_mzml.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,153 @@
"""Generate a minimal synthetic DIA mzML file for CI testing.

The file contains 5 DIA cycles (5 MS1 + 15 MS2 = 20 spectra total) with
realistic isolation windows and small random peak arrays.
"""
import base64
import struct
from pathlib import Path

import numpy as np

RNG = np.random.default_rng(42)

# DIA isolation windows (center, half-width) in m/z
WINDOWS = [(412.5, 12.5), (437.5, 12.5), (462.5, 12.5)]
N_CYCLES = 5
CYCLE_INTERVAL_MIN = 0.5 # 30 s per cycle
N_MS1_PEAKS = 60
N_MS2_PEAKS = 40


def _encode_float32(arr: np.ndarray) -> str:
data = struct.pack(f"{len(arr)}f", *arr.astype(np.float64))
return base64.b64encode(data).decode("ascii")


def _binary_array(arr: np.ndarray, accession: str, name: str) -> str:
encoded = _encode_float32(arr)
return (
f' <binaryDataArray encodedLength="{len(encoded)}">\n'
f' <cvParam cvRef="MS" accession="{accession}" name="{name}"/>\n'
f' <cvParam cvRef="MS" accession="MS:1000521" name="32-bit float"/>\n'
f' <cvParam cvRef="MS" accession="MS:1000576" name="no compression"/>\n'
f" <binary>{encoded}</binary>\n"
f" </binaryDataArray>"
)


def _ms1_spectrum(idx: int, rt_min: float) -> str:
mz = np.sort(RNG.uniform(300, 1200, N_MS1_PEAKS)).astype(np.float32)
intensity = RNG.exponential(1e6, N_MS1_PEAKS).astype(np.float32)
mz_block = _binary_array(mz, "MS:1000514", "m/z array")
int_block = _binary_array(intensity, "MS:1000515", "intensity array")
return f"""\
<spectrum index="{idx}" id="scan={idx + 1}" defaultArrayLength="{N_MS1_PEAKS}">
<cvParam cvRef="MS" accession="MS:1000511" name="ms level" value="1"/>
<scanList count="1">
<scan>
<cvParam cvRef="MS" accession="MS:1000016" name="scan start time"
value="{rt_min:.4f}" unitCvRef="UO" unitAccession="UO:0000031" unitName="minute"/>
</scan>
</scanList>
<binaryDataArrayList count="2">
{mz_block}
{int_block}
</binaryDataArrayList>
</spectrum>"""


def _ms2_spectrum(
idx: int, rt_min: float, center: float, lower: float, upper: float
) -> str:
mz = np.sort(RNG.uniform(100, 1500, N_MS2_PEAKS)).astype(np.float32)
intensity = RNG.exponential(1e5, N_MS2_PEAKS).astype(np.float32)
mz_block = _binary_array(mz, "MS:1000514", "m/z array")
int_block = _binary_array(intensity, "MS:1000515", "intensity array")
return f"""\
<spectrum index="{idx}" id="scan={idx + 1}" defaultArrayLength="{N_MS2_PEAKS}">
<cvParam cvRef="MS" accession="MS:1000511" name="ms level" value="2"/>
<scanList count="1">
<scan>
<cvParam cvRef="MS" accession="MS:1000016" name="scan start time"
value="{rt_min:.4f}" unitCvRef="UO" unitAccession="UO:0000031" unitName="minute"/>
</scan>
</scanList>
<precursorList count="1">
<precursor>
<isolationWindow>
<cvParam cvRef="MS" accession="MS:1000827" name="isolation window target m/z" value="{center}"/>
<cvParam cvRef="MS" accession="MS:1000828" name="isolation window lower offset" value="{lower}"/>
<cvParam cvRef="MS" accession="MS:1000829" name="isolation window upper offset" value="{upper}"/>
</isolationWindow>
</precursor>
</precursorList>
<binaryDataArrayList count="2">
{mz_block}
{int_block}
</binaryDataArrayList>
</spectrum>"""


def create_mini_mzml(output_path: str | Path) -> Path:
"""Write a minimal DIA mzML to *output_path* and return the path."""
output_path = Path(output_path)
spectra = []
idx = 0
for cycle in range(N_CYCLES):
rt_ms1 = (cycle + 1) * CYCLE_INTERVAL_MIN
spectra.append(_ms1_spectrum(idx, rt_ms1))
idx += 1
for win_idx, (center, hw) in enumerate(WINDOWS):
rt_ms2 = rt_ms1 + (win_idx + 1) * 0.005
spectra.append(_ms2_spectrum(idx, rt_ms2, center, hw, hw))
idx += 1

n_total = len(spectra)
spectrum_block = "\n".join(spectra)

xml = f"""\
<?xml version="1.0" encoding="utf-8"?>
<mzML xmlns="http://psi.hupo.org/ms/mzml"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://psi.hupo.org/ms/mzml http://psidev.info/files/ms/mzML/xsd/mzML1.1.0.xsd">
<cvList count="2">
<cv id="MS" fullName="Proteomics Standards Initiative Mass Spectrometry Ontology"
URI="https://raw.githubusercontent.com/HUPO-PSI/psi-ms-CV/master/psi-ms.obo"/>
<cv id="UO" fullName="Unit Ontology"
URI="https://raw.githubusercontent.com/bio-ontology-research-group/unit-ontology/master/unit.obo"/>
</cvList>
<fileDescription>
<fileContent>
<cvParam cvRef="MS" accession="MS:1000580" name="MSn spectrum" value=""/>
</fileContent>
</fileDescription>
<softwareList count="1">
<software id="test_gen" version="1.0">
<cvParam cvRef="MS" accession="MS:1000799" name="custom unreleased software tool" value=""/>
</software>
</softwareList>
<dataProcessingList count="1">
<dataProcessing id="dp">
<processingMethod order="0" softwareRef="test_gen">
<cvParam cvRef="MS" accession="MS:1000544" name="Conversion to mzML" value=""/>
</processingMethod>
</dataProcessing>
</dataProcessingList>
<run>
<spectrumList count="{n_total}" defaultDataProcessingRef="dp">
{spectrum_block}
</spectrumList>
</run>
</mzML>
"""
output_path.write_text(xml, encoding="utf-8")
return output_path


if __name__ == "__main__":
import sys

out = Path(sys.argv[1]) if len(sys.argv) > 1 else Path("mini_demo.mzML")
p = create_mini_mzml(out)
print(f"Wrote {p} ({p.stat().st_size} bytes)")
44 changes: 44 additions & 0 deletions tests/test_e2e.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
"""End-to-end test: run ``cascadia sequence`` with a dummy checkpoint.

Requires the dummy checkpoint to exist at the path given by the
CASCADIA_DUMMY_CKPT environment variable (set by the CI workflow).
Skipped automatically when the variable is absent (e.g. local dev without
the checkpoint).
"""
import os
import sys
from pathlib import Path

import pytest

TESTS_DIR = Path(__file__).parent
CKPT_PATH = str(Path(os.environ["CASCADIA_DUMMY_CKPT"]).expanduser()) if os.environ.get("CASCADIA_DUMMY_CKPT") else ""


@pytest.mark.skipif(not CKPT_PATH, reason="CASCADIA_DUMMY_CKPT not set")
def test_sequence_command(tmp_path):
"""Full ``cascadia sequence`` pipeline on the mini mzML with a dummy checkpoint."""
sys.path.insert(0, str(TESTS_DIR))
from create_mini_mzml import create_mini_mzml

mzml_file = create_mini_mzml(tmp_path / "mini_demo.mzML")
out_prefix = str(tmp_path / "ci_results")

# Drive the CLI entry-point directly so we test the real code path.
sys.argv = [
"cascadia",
"sequence",
str(mzml_file),
CKPT_PATH,
"--out", out_prefix,
"--batch_size", "4",
"--max_charge", "2",
]

from cascadia.cascadia import sequence
sequence()

ssl_file = Path(out_prefix + ".ssl")
assert ssl_file.exists(), f"Expected output file {ssl_file} was not created"
content = ssl_file.read_text()
assert len(content) > 0, "Output SSL file is empty"
Loading
Loading