G-CoS: An Interpretable Gain-Cost Framework for User Satisfaction Estimation in Generative Information Retrieval
Official code repository for the SIGIR 2026 short paper "G-CoS: An Interpretable Gain-Cost Framework for User Satisfaction Estimation in Generative Information Retrieval".
Interpretable satisfaction estimation for generative information retrieval.
G-CoS is an interpretable framework for user satisfaction estimation in Generative Information Retrieval (GenIR). The repository includes the main G-CoS implementation, feature-based and sequence baselines, LLM-based baselines, and lightweight utilities for local smoke testing.
- Interpretable gain-cost formulation for GenIR satisfaction estimation
- Baseline implementations for Ridge, Random Forest, XGBoost, MLP, LSTM, Transformer, SPUR, G-Eval, URS, and GAH
- Synthetic smoke-test workflow for running the repository without the original dataset
- Public data schema documentation for adapting the code to new datasets
G-CoS/
├── src/models/ # G-CoS implementation
├── src/baselines/ # Feature-based and sequence baselines
├── src/components/ # Semantic gain and frustration estimation
├── src/llm_baselines/ # G-Eval, URS, GAH, and SPUR baselines
├── scripts/ # Synthetic smoke-test data generation
├── data/ # Data schema documentation
├── requirements.txt
└── README.md
- Python 3.8+
- PyTorch 1.12+ for LSTM/Transformer baselines
- See
requirements.txtfor the full dependency list
cd G-CoS
python3 -m venv venv
source venv/bin/activate # On macOS/Linux
# venv\Scripts\activate # On Windows
pip install -r requirements.txt
# Optional: prepare local API credentials
cp .env.example .envThe commands below assume you already have a behavioral CSV and a semantic score JSON.
mkdir -p outputs/gcos_results
python3 src/models/gcos_model.py \
--data data/behavior_sessions.csv \
--semantic data/semantic_scores.json \
--output outputs/gcos_results \
--seed 2024python3 src/baselines/traditional_ml_baselines.py \
--data data/behavior_sessions.csv \
--semantic data/semantic_scores.json \
--output outputs/baseline_results \
--seed 2024Frustration Cost (F):
python3 src/components/frustration_cost_estimation.py \
--data data/behavior_sessions.csv \
--output_json outputs/frustration_cost_scores.json \
--output_csv data/behavior_sessions_with_frustration.csvThe script also accepts the legacy aliases --csv and --output for backward compatibility.
Semantic Gain (G):
export DEEPSEEK_API_KEY="your-api-key"
python3 src/components/semantic_gain_estimation.py \
--data data/behavior_sessions.csv \
--output outputs/semantic_gain_scores.jsonpython3 src/llm_baselines/spur_baseline.py \
--data data/behavior_sessions.csv \
--output outputs/spur_scores.json
python3 src/llm_baselines/llm_judge_baselines.py \
--data data/behavior_sessions.csv \
--methods geval,urs,gah \
--output outputs/llm_judge_results.jsonTo run only a subset of methods:
python3 src/llm_baselines/llm_judge_baselines.py \
--data data/behavior_sessions.csv \
--methods geval,gah \
--output outputs/llm_judge_results.jsonIf you just want to verify that the repository runs end-to-end without access to the original dataset, use the synthetic demo:
python3 scripts/generate_synthetic_data.py \
--samples 36 \
--seed 2024 \
--data-output data/behavior_sessions.csv \
--semantic-output data/semantic_scores.json
python3 src/models/gcos_model.py \
--data data/behavior_sessions.csv \
--semantic data/semantic_scores.json \
--output outputs/gcos_results \
--seed 2024 \
--folds 3
python3 src/baselines/traditional_ml_baselines.py \
--data data/behavior_sessions.csv \
--semantic data/semantic_scores.json \
--output outputs/baseline_results \
--seed 2024 \
--folds 3 \
--n-resamples 200The synthetic workflow is for smoke testing only and does not reproduce the paper's reported metrics.
The main scripts require two input files: one behavioral interaction CSV and one semantic score JSON. The filenames below are only examples; you can place them anywhere and pass their paths with --data and --semantic.
-
Behavioral data CSV (for example,
data/behavior_sessions.csv)- Required columns include
satisfaction,total_duration,click_count,query_reform_count,qa_pair_text,data_source - Required task metadata includes either
task_idortask_type - Optional columns include
frustration_weight
- Required columns include
-
Semantic scores JSON (for example,
data/semantic_scores.json)- LLM-generated semantic quality scores
- Format:
[{"index": 0, "geval_score": 4.5, "satisfaction_true": 5}, ...]
Full schema details are documented in data/DATA.md.
The experiments in this repository are based on the dataset introduced in:
Liang, Y., Wu, Z., Zhang, F., Song, D., and Huang, H. (2025). "How Users Interact with Generative Information Retrieval Systems: A Study of User Behavior and Search Experience." In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '25), 634-644. https://doi.org/10.1145/3726302.3729998
This repository does not redistribute the original dataset. Please obtain the data from the original source and follow the data access, licensing, and usage terms specified by the original authors or publisher.
If you do not have access to the original data, you can still run the synthetic smoke test provided in scripts/generate_synthetic_data.py, or adapt the code to your own dataset by following the schema documented in data/DATA.md.
Several scripts require API access for LLM-based evaluation:
- DeepSeek API: set
DEEPSEEK_API_KEY - OpenAI API: set
OPENAI_API_KEY
If you use this code in your research, please cite:
@inproceedings{shi2026gcos,
title={G-CoS: An Interpretable Gain-Cost Framework for User Satisfaction Estimation in Generative Information Retrieval},
author={Shi, Jia-Ling and Wu, Zhijing and Liang, Yidong and Mao, Xian-Ling},
booktitle={Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval},
year={2026},
doi={10.1145/3805712.3809934}
}This code is released under the Apache License 2.0. See LICENSE for details.
For questions or issues, please open a GitHub issue or contact Jia-Ling Shi at sjl@bit.edu.cn.