Skip to content
Academic-HammerPublic

About

G-CoS: An interpretable gain-cost framework for modeling user satisfaction via response quality and interaction cost.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

G-CoS: An Interpretable Gain-Cost Framework for User Satisfaction Estimation in Generative Information Retrieval

Official code repository for the SIGIR 2026 short paper "G-CoS: An Interpretable Gain-Cost Framework for User Satisfaction Estimation in Generative Information Retrieval".

Interpretable satisfaction estimation for generative information retrieval.

Overview

G-CoS is an interpretable framework for user satisfaction estimation in Generative Information Retrieval (GenIR). The repository includes the main G-CoS implementation, feature-based and sequence baselines, LLM-based baselines, and lightweight utilities for local smoke testing.

Highlights

  • Interpretable gain-cost formulation for GenIR satisfaction estimation
  • Baseline implementations for Ridge, Random Forest, XGBoost, MLP, LSTM, Transformer, SPUR, G-Eval, URS, and GAH
  • Synthetic smoke-test workflow for running the repository without the original dataset
  • Public data schema documentation for adapting the code to new datasets

Repository Structure

G-CoS/
├── src/models/          # G-CoS implementation
├── src/baselines/       # Feature-based and sequence baselines
├── src/components/      # Semantic gain and frustration estimation
├── src/llm_baselines/   # G-Eval, URS, GAH, and SPUR baselines
├── scripts/             # Synthetic smoke-test data generation
├── data/                # Data schema documentation
├── requirements.txt
└── README.md

Requirements

  • Python 3.8+
  • PyTorch 1.12+ for LSTM/Transformer baselines
  • See requirements.txt for the full dependency list

Installation

cd G-CoS

python3 -m venv venv
source venv/bin/activate  # On macOS/Linux
# venv\Scripts\activate   # On Windows

pip install -r requirements.txt

# Optional: prepare local API credentials
cp .env.example .env

Quick Start

The commands below assume you already have a behavioral CSV and a semantic score JSON.

1. Train G-CoS

mkdir -p outputs/gcos_results

python3 src/models/gcos_model.py \
  --data data/behavior_sessions.csv \
  --semantic data/semantic_scores.json \
  --output outputs/gcos_results \
  --seed 2024

2. Run Baseline Experiments

python3 src/baselines/traditional_ml_baselines.py \
  --data data/behavior_sessions.csv \
  --semantic data/semantic_scores.json \
  --output outputs/baseline_results \
  --seed 2024

3. Generate Component Scores

Frustration Cost (F):

python3 src/components/frustration_cost_estimation.py \
  --data data/behavior_sessions.csv \
  --output_json outputs/frustration_cost_scores.json \
  --output_csv data/behavior_sessions_with_frustration.csv

The script also accepts the legacy aliases --csv and --output for backward compatibility.

Semantic Gain (G):

export DEEPSEEK_API_KEY="your-api-key"

python3 src/components/semantic_gain_estimation.py \
  --data data/behavior_sessions.csv \
  --output outputs/semantic_gain_scores.json

4. Run LLM-as-a-Judge Baselines

python3 src/llm_baselines/spur_baseline.py \
  --data data/behavior_sessions.csv \
  --output outputs/spur_scores.json

python3 src/llm_baselines/llm_judge_baselines.py \
  --data data/behavior_sessions.csv \
  --methods geval,urs,gah \
  --output outputs/llm_judge_results.json

To run only a subset of methods:

python3 src/llm_baselines/llm_judge_baselines.py \
  --data data/behavior_sessions.csv \
  --methods geval,gah \
  --output outputs/llm_judge_results.json

Smoke Test

If you just want to verify that the repository runs end-to-end without access to the original dataset, use the synthetic demo:

python3 scripts/generate_synthetic_data.py \
  --samples 36 \
  --seed 2024 \
  --data-output data/behavior_sessions.csv \
  --semantic-output data/semantic_scores.json

python3 src/models/gcos_model.py \
  --data data/behavior_sessions.csv \
  --semantic data/semantic_scores.json \
  --output outputs/gcos_results \
  --seed 2024 \
  --folds 3

python3 src/baselines/traditional_ml_baselines.py \
  --data data/behavior_sessions.csv \
  --semantic data/semantic_scores.json \
  --output outputs/baseline_results \
  --seed 2024 \
  --folds 3 \
  --n-resamples 200

The synthetic workflow is for smoke testing only and does not reproduce the paper's reported metrics.

Data Preparation

The main scripts require two input files: one behavioral interaction CSV and one semantic score JSON. The filenames below are only examples; you can place them anywhere and pass their paths with --data and --semantic.

  1. Behavioral data CSV (for example, data/behavior_sessions.csv)

    • Required columns include satisfaction, total_duration, click_count, query_reform_count, qa_pair_text, data_source
    • Required task metadata includes either task_id or task_type
    • Optional columns include frustration_weight
  2. Semantic scores JSON (for example, data/semantic_scores.json)

    • LLM-generated semantic quality scores
    • Format: [{"index": 0, "geval_score": 4.5, "satisfaction_true": 5}, ...]

Full schema details are documented in data/DATA.md.

Data Availability

The experiments in this repository are based on the dataset introduced in:

Liang, Y., Wu, Z., Zhang, F., Song, D., and Huang, H. (2025). "How Users Interact with Generative Information Retrieval Systems: A Study of User Behavior and Search Experience." In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '25), 634-644. https://doi.org/10.1145/3726302.3729998

This repository does not redistribute the original dataset. Please obtain the data from the original source and follow the data access, licensing, and usage terms specified by the original authors or publisher.

If you do not have access to the original data, you can still run the synthetic smoke test provided in scripts/generate_synthetic_data.py, or adapt the code to your own dataset by following the schema documented in data/DATA.md.

API Keys

Several scripts require API access for LLM-based evaluation:

  • DeepSeek API: set DEEPSEEK_API_KEY
  • OpenAI API: set OPENAI_API_KEY

Citation

If you use this code in your research, please cite:

@inproceedings{shi2026gcos,
  title={G-CoS: An Interpretable Gain-Cost Framework for User Satisfaction Estimation in Generative Information Retrieval},
  author={Shi, Jia-Ling and Wu, Zhijing and Liang, Yidong and Mao, Xian-Ling},
  booktitle={Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval},
  year={2026},
  doi={10.1145/3805712.3809934}
}

License

This code is released under the Apache License 2.0. See LICENSE for details.

Contact

For questions or issues, please open a GitHub issue or contact Jia-Ling Shi at sjl@bit.edu.cn.

About

G-CoS: An interpretable gain-cost framework for modeling user satisfaction via response quality and interaction cost.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages