Code for training and evaluating gene and protein disambiguation models on the CoDiet and JNLPBA corpora, including both within-corpus and cross-corpus evaluation.
.
├── data
│ ├── original
│ └── processed
├── dataset.py
├── evaluation.py
├── exp
│ ├── CoDiet
│ └── JNLPBA
├── jnlpba
│ ├── evaluate_j.py
│ ├── preprocess_j.py
│ ├── split_j.py
│ └── train_j.py
├── metrics.py
├── multi_training_bert.py
├── prediction.py
├── preprocess.py
├── README.md
├── relaxed_f1_per_class.py
├── requirements.txt
└── utils.py
git clone https://github.com/omicsNLP/geneProt_Disambiguation.git
cd geneProt_Disambiguationconda create -n gene-protein-dis python=3.12
conda activate gene-protein-dis
pip install -r requirements.txtUsed for sentence-boundary detection when splitting long passages.
python -m spacy download en_core_web_smCreate a data folder at the project root, then download and extract the CoDiet datasets into it.
mkdir -p data/original
cd data/original
# CoDiet-Gold-Public set (used for training/evaluation)
wget https://zenodo.org/records/17610205/files/CoDiet-Gold-public.zip
unzip CoDiet-Gold-public.zip
# CoDiet-Gold-Private set (used only for prediction and predicted on codabench)
wget https://zenodo.org/records/17610205/files/CoDiet-Gold-private.zip
unzip CoDiet-Gold-private.zipAfter extraction you should have:
data/original/CoDiet-Gold-public/*.json (450 files)
data/original/CoDiet-Gold-private/*.json (50 files)
For JNLPBA, download the original data from the official GENIA project source and then extract it:
wget http://www.nactem.ac.uk/GENIA/current/Shared-tasks/JNLPBA/Train/Genia4ERtraining.tar.gz
wget http://www.nactem.ac.uk/GENIA/current/Shared-tasks/JNLPBA/Evaluation/Genia4ERtest.tar.gz
mkdir -p GENIA_train/
tar -xzf Genia4ERtraining.tar.gz -C GENIA_train
mkdir -p GENIA_test/
tar -xzf Genia4ERtest.tar.gz -C GENIA_test
cd ../..After extraction, split the raw data into train/validation/test set in train.tsv, devel.tsv, and test.tsv under ./data/original/JNLPBA/:
python jnlpba/split_j.py
Preprocessed/tokenised data and trained model checkpoints will be under ./data/processed/ and ./exp/ respectively, both created automatically.
--model_name accepts a short name from the table below (results reported in the paper), or or any full HuggingFace model path (e.g., org/model) — the pipeline isn't limited to the models we evaluated, feel free to try others. When using a full path, / is replaced with _ in output directory names.
| Model name | HuggingFace model path |
|---|---|
| biobert-base | dmis-lab/biobert-base-cased-v1.1 |
| biobert-large | dmis-lab/biobert-large-cased-v1.1-squad |
| BioLinkBERT-base | michiyasunaga/BioLinkBERT-base |
| BioLinkBERT-large | michiyasunaga/BioLinkBERT-large |
| biomedbert-base | microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext |
| Bio_ClinicalBERT | emilyalsentzer/Bio_ClinicalBERT |
| Clinical-Longformer | yikuan8/Clinical-Longformer |
multi_training_bert.py (CoDiet) and jnlpba/train_j.py (JNLPBA) share the same argument set. All arguments are optional (have defaults), there are no required arguments. The only difference between these two is the
default value of --corpus_name
| Argument | Type | Default | Description |
|---|---|---|---|
model_name |
str | biobert-base | Short name from the supported model table, or a full HuggingFace path |
num_epochs |
int | 10 | Training epochs |
batch_size |
int | 32 | Batch size |
max_tokens |
int | 256 | Max token length |
corpus_name |
str | CoDiet (multi_training_bert.py) or JNLPBA (train_j.py) |
Corpus used for training and output directory |
trained_on |
str | same as --corpus_name | Corpus the evaluated model was trained on (for cross-corpus eval) |
lr |
float | 3e-5 | Learning rate |
train |
flag | off | Run training (otherwise only evaluation) |
python multi_training_bert.py \
--train \
--model_name biobert-base \
--num_epochs 10 \
--batch_size 32 \
--max_tokens 256You're not limited to the models in the table above. Any HuggingFace model path works, e.g.,
python multi_training_bert.py \
--train \
--model_name FacebookAI/roberta-base \
--num_epochs 10 \
--batch_size 32 \
--max_tokens 256Drop --train to skip training and only run evaluation using an existing best_model.pt in the corresponding ./exp/{corpus_name}/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/ directory.
python multi_training_bert.py \
--model_name biobert-base \
--num_epochs 10 \
--batch_size 32 \
--max_tokens 256Prerequisite: both corpora must already have a trained model under
./exp/{corpus_name}/ before running cross-corpus evaluation. If best_model.pt doesn't exist, the script prints Best model not found and uses initialised (untrained) model, which produces meaningless results.
To evaluate a model trained on one corpus against another corpus's test set, set --corpus_name to the corpus being evaluated on and --trained_on to the corpus the model was trained on.
--model_name, --num_epochs, --batch_size, and --max_tokens must match the original settings used for training in order to build the path to best_model.pt under ./exp/{trained_on}/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/:
python multi_training_bert.py \
--corpus_name CoDiet \
--trained_on JNLPBA \
--model_name biobert-base \
--num_epochs 10 \
--batch_size 32 \
--max_tokens 256This loads the model from ./exp/JNLPBA/ and evaluates it on the CoDiet test set. Don't set --train here since cross-corpus scripts only evaluate, they don't train.
Cross-corpus evaluation results are saved separately from the model itself, under ./exp/{corpus_name}/{trained_on}_trained/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/
Example: evaluating a JNLPBA-trained model on CoDiet, their results will be saved to
./exp/CoDiet/JNLPBA_trained/, while the best model is under./exp/JNLPBA/
python jnlpba/train_j.py \
--train \
--model_name biobert-base \
--num_epochs 10 \
--batch_size 32 \
--max_tokens 256Cross-corpus evaluation works the same, e.g., evaluating a CoDiet-trained model on JNLPBA:
python jnlpba/train_j.py \
--corpus_name JNLPBA \
--trained_on CoDiet \
--model_name biobert-base \
--num_epochs 10 \
--batch_size 32 \
--max_tokens 256Cross-corpus evaluation results are saved in ./exp/JNLPBA/CoDiet_trained/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/
prediction.py runs a trained model on unannotated CoDiet Gold private subset and generates a .zip file including all predictions, ready to be uploaded directly to Codabench for evaluation.
python prediction.py \
--model_dir ./exp/CoDiet/biobert-base_e10_bs32_256 --model_dir(required): a trained experiment directory containingbest_model.pt,label_mappings.json, andtrain_config.json.--input_dir: defaults to./data/original/CoDiet-Gold-private.--output_dir: defaults to{model_dir}/predictions.
Output will be saved to --output_dir:
pred_*.jsonfor each input file, with per-passage entity predictions.all_predictions.csv, a table of every predicted entity across all files.{output_dir_name}_predictions.zip, the CSV above zipped for direct uploading.
Only GENE and PROTEIN entities are extracted. Any other annotation type present in the raw data is removed during preprocessing.