This project is for my First Year Research Experience at the University of Oklahoma. I am doing this project under the supervision of Data Institute for Societal Challenges faculty Dr. Wolfgang.
The goal is to create a "search engine" for social media posts. To do this I created an algorithm based upon active feedback through uncertainty sampling.
Want to run it? Jump to Quickstart.
Presented at the University of Oklahoma First-Year Research Experience (FYRE). Click for full resolution.
Problem: Classifying social media posts is hard because of sarcasm, slang, and ambiguous language.
Objective: Build a model that labels social media content as relevant or not, using active learning and contextual understanding.
Sentiment140: 1.6 million tweets from 2009, collected with the Twitter API.
Two-stage learning:
- Initial training. Tweets are preprocessed (tokenization, vectorization) and used to train three baselines:
- Logistic regression + bag of words (single-word frequency)
- Logistic regression + bi-grams (word-pair frequency)
- BERT (contextual understanding)
- Active feedback loop. Using uncertainty sampling, the model finds the post it is least sure about and shows it to a human. The human labels it, and the model updates. Applied to BERT, this is AL-BERT (active-learning BERT).
Approximate accuracy, read from the poster's charts:
| Model | Without feedback (1,000 prelabeled) | With active feedback (100 inputs) |
|---|---|---|
| Bag of words | ~0.76 | ~0.77 |
| Bi-grams | ~0.78 | ~0.83 |
| BERT / AL-BERT | ~0.89 | ~0.95 |
With 100 actively chosen labels, every model matched or beat the same model trained on 1,000 traditional labels. AL-BERT did best.
- Active feedback works. Models that got human feedback beat the traditional method, because the feedback resolved slang and sarcasm.
- Context is key. AL-BERT beat the simpler models because it understands whole sentences, not just individual words.
- Challenges. The 2009 dataset is missing modern slang, and the loop needs human effort, though that effort is what lets the model adapt.
- Future steps. Use newer data, and cut human effort with automated feedback.
Combining a model's ability to process millions of tweets with human insight into nuance produces a more accurate classifier for ambiguous domains such as healthcare, education, customer service, legal, social media, and news.
Requires Python 3.11 (TensorFlow 2.15 doesn't support newer versions). A GPU is optional; everything runs on CPU.
git clone https://github.com/dishishshawn/Relevance-Classification-Model.git
cd Relevance-Classification-Model
uv venv -p 3.11 && source .venv/bin/activate # or: python3.11 -m venv .venv
uv pip install -r requirements.txt # or: pip install -r requirements.txtThen run the pipeline in order. Everything is written to data/ (git-ignored), and the first command downloads Sentiment140 (~80 MB) for you.
# 1. Label tweets by keyword: a tweet is "relevant" if it mentions the topic
python prepare_data.py test football # held-out test set -> data/test
python prepare_data.py train football # balanced sets of 25, 50, ... per class -> data/25.25, data/50.50, ...
# 2. Baseline without feedback: accuracy vs. training-set size
python baseline.py # -> results/baseline_sweep.csv
# 3. Active feedback: each script shows you tweets and asks "Relevant? (y/n/stop)"
python active_bow.py # bag of words + uncertainty sampling
python active_lstm.py # LSTM
python active_bert.py # AL-BERT (downloads bert-base-uncased, ~440 MB)
# 4. Compare the feedback-trained models on the test set
python evaluate.py # -> results/model_comparison.csvEvery script takes --help. The feedback scripts save their progress to data/feedback_*/, so you can type stop and resume later. Any keyword works in place of football, and you can pass several (python prepare_data.py test football soccer nfl).
| File | Purpose |
|---|---|
common.py |
Paths, Sentiment140 download and loading, text cleaning (URLs, punctuation, stopwords) |
prepare_data.py |
Builds the keyword-labeled training sets and test set |
baseline.py |
TF-IDF + logistic regression trained on each set size, scored on the test set |
active_bow.py |
Bag of words with modAL uncertainty sampling |
active_lstm.py |
Keras LSTM with a relevance feedback loop |
active_bert.py |
AL-BERT: fine-tuned bert-base-uncased, scoring each tweet by uncertainty + relevance |
evaluate.py |
Scores every trained feedback model on data/test |
results/baseline_sweep.csv |
Baseline results from the original research run |
poster/ |
FYRE poster |
This work is supported in part by the Data Institute for Societal Challenges, University of Oklahoma.
- Cook, T. (2020, August 30). How BERT determines search relevance. Towards Data Science.
- Sentiment140.
- Baumgärtner, T., Ribeiro, L. F. R., Reimers, N., & Gurevych, I. (2022). Incorporating relevance feedback for information-seeking retrieval using few-shot document re-ranking. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics.
