Skip to content

Repository files navigation

Gene and Protein Entity Disambiguation

Code for training and evaluating gene and protein disambiguation models on the CoDiet and JNLPBA corpora, including both within-corpus and cross-corpus evaluation.

Repo tree

.
├── data
│   ├── original
│   └── processed
├── dataset.py
├── evaluation.py
├── exp
│   ├── CoDiet
│   └── JNLPBA
├── jnlpba
│   ├── evaluate_j.py
│   ├── preprocess_j.py
│   ├── split_j.py
│   └── train_j.py
├── metrics.py
├── multi_training_bert.py
├── prediction.py
├── preprocess.py
├── README.md
├── relaxed_f1_per_class.py
├── requirements.txt
└── utils.py

Setup

1. Clone this repo

git clone https://github.com/omicsNLP/geneProt_Disambiguation.git
cd geneProt_Disambiguation

2. Create a virtual environment

conda create -n gene-protein-dis python=3.12
conda activate gene-protein-dis
pip install -r requirements.txt

3. Download the spaCy model

Used for sentence-boundary detection when splitting long passages.

python -m spacy download en_core_web_sm

Data

Create a data folder at the project root, then download and extract the CoDiet datasets into it.

mkdir -p data/original
cd data/original

# CoDiet-Gold-Public set (used for training/evaluation)
wget https://zenodo.org/records/17610205/files/CoDiet-Gold-public.zip
unzip CoDiet-Gold-public.zip

# CoDiet-Gold-Private set (used only for prediction and predicted on codabench)
wget https://zenodo.org/records/17610205/files/CoDiet-Gold-private.zip
unzip CoDiet-Gold-private.zip

After extraction you should have:

data/original/CoDiet-Gold-public/*.json (450 files)
data/original/CoDiet-Gold-private/*.json (50 files)

For JNLPBA, download the original data from the official GENIA project source and then extract it:

wget http://www.nactem.ac.uk/GENIA/current/Shared-tasks/JNLPBA/Train/Genia4ERtraining.tar.gz

wget http://www.nactem.ac.uk/GENIA/current/Shared-tasks/JNLPBA/Evaluation/Genia4ERtest.tar.gz

mkdir -p GENIA_train/
tar -xzf Genia4ERtraining.tar.gz -C GENIA_train

mkdir -p GENIA_test/
tar -xzf Genia4ERtest.tar.gz -C GENIA_test

cd ../..

After extraction, split the raw data into train/validation/test set in train.tsv, devel.tsv, and test.tsv under ./data/original/JNLPBA/:

python jnlpba/split_j.py

Preprocessed/tokenised data and trained model checkpoints will be under ./data/processed/ and ./exp/ respectively, both created automatically.

Supported models

--model_name accepts a short name from the table below (results reported in the paper), or or any full HuggingFace model path (e.g., org/model) — the pipeline isn't limited to the models we evaluated, feel free to try others. When using a full path, / is replaced with _ in output directory names.

Model name HuggingFace model path
biobert-base dmis-lab/biobert-base-cased-v1.1
biobert-large dmis-lab/biobert-large-cased-v1.1-squad
BioLinkBERT-base michiyasunaga/BioLinkBERT-base
BioLinkBERT-large michiyasunaga/BioLinkBERT-large
biomedbert-base microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext
Bio_ClinicalBERT emilyalsentzer/Bio_ClinicalBERT
Clinical-Longformer yikuan8/Clinical-Longformer

Arguments

multi_training_bert.py (CoDiet) and jnlpba/train_j.py (JNLPBA) share the same argument set. All arguments are optional (have defaults), there are no required arguments. The only difference between these two is the default value of --corpus_name

Argument Type Default Description
model_name str biobert-base Short name from the supported model table, or a full HuggingFace path
num_epochs int 10 Training epochs
batch_size int 32 Batch size
max_tokens int 256 Max token length
corpus_name str CoDiet (multi_training_bert.py) or JNLPBA (train_j.py) Corpus used for training and output directory
trained_on str same as --corpus_name Corpus the evaluated model was trained on (for cross-corpus eval)
lr float 3e-5 Learning rate
train flag off Run training (otherwise only evaluation)

Training and evaluation (CoDiet)

python multi_training_bert.py \
  --train \
  --model_name biobert-base \
  --num_epochs 10 \
  --batch_size 32 \
  --max_tokens 256

You're not limited to the models in the table above. Any HuggingFace model path works, e.g.,

python multi_training_bert.py \
	--train \
	--model_name FacebookAI/roberta-base \
	--num_epochs 10 \
	--batch_size 32 \
	--max_tokens 256

Drop --train to skip training and only run evaluation using an existing best_model.pt in the corresponding ./exp/{corpus_name}/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/ directory.

python multi_training_bert.py \
  --model_name biobert-base \
  --num_epochs 10 \
  --batch_size 32 \
  --max_tokens 256

Cross-corpus evaluation

Prerequisite: both corpora must already have a trained model under ./exp/{corpus_name}/ before running cross-corpus evaluation. If best_model.pt doesn't exist, the script prints Best model not found and uses initialised (untrained) model, which produces meaningless results.

To evaluate a model trained on one corpus against another corpus's test set, set --corpus_name to the corpus being evaluated on and --trained_on to the corpus the model was trained on. --model_name, --num_epochs, --batch_size, and --max_tokens must match the original settings used for training in order to build the path to best_model.pt under ./exp/{trained_on}/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/:

python multi_training_bert.py \
  --corpus_name CoDiet \
  --trained_on JNLPBA \
  --model_name biobert-base \
  --num_epochs 10 \
  --batch_size 32 \
  --max_tokens 256

This loads the model from ./exp/JNLPBA/ and evaluates it on the CoDiet test set. Don't set --train here since cross-corpus scripts only evaluate, they don't train.

Cross-corpus evaluation results are saved separately from the model itself, under ./exp/{corpus_name}/{trained_on}_trained/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/

Example: evaluating a JNLPBA-trained model on CoDiet, their results will be saved to ./exp/CoDiet/JNLPBA_trained/, while the best model is under ./exp/JNLPBA/

Training and evaluation (JNLPBA)

python jnlpba/train_j.py \
  --train \
  --model_name biobert-base \
  --num_epochs 10 \
  --batch_size 32 \
  --max_tokens 256

Cross-corpus evaluation works the same, e.g., evaluating a CoDiet-trained model on JNLPBA:

python jnlpba/train_j.py \
  --corpus_name JNLPBA \
  --trained_on CoDiet \
  --model_name biobert-base \
  --num_epochs 10 \
  --batch_size 32 \
  --max_tokens 256

Cross-corpus evaluation results are saved in ./exp/JNLPBA/CoDiet_trained/{model_name}_e{epochs}_bs{batch_size}_{max_tokens}/

Prediction on unannotated documents (CoDiet-Gold-private)

prediction.py runs a trained model on unannotated CoDiet Gold private subset and generates a .zip file including all predictions, ready to be uploaded directly to Codabench for evaluation.

python prediction.py \
  --model_dir ./exp/CoDiet/biobert-base_e10_bs32_256 
  • --model_dir (required): a trained experiment directory containing best_model.pt, label_mappings.json, and train_config.json.
  • --input_dir: defaults to ./data/original/CoDiet-Gold-private.
  • --output_dir: defaults to {model_dir}/predictions.

Output will be saved to --output_dir:

  • pred_*.json for each input file, with per-passage entity predictions.
  • all_predictions.csv, a table of every predicted entity across all files.
  • {output_dir_name}_predictions.zip, the CSV above zipped for direct uploading.

Notes

Only GENE and PROTEIN entities are extracted. Any other annotation type present in the raw data is removed during preprocessing.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages