Longitudinal human-LLM interaction study documenting emergent self-descriptive frameworks in a GPT-5.4 instance across 23 days and comparative cold sessions.
-
Updated
Aug 15, 2026 - Python
Longitudinal human-LLM interaction study documenting emergent self-descriptive frameworks in a GPT-5.4 instance across 23 days and comparative cold sessions.
Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
Do language models show non-verbal signs of adverse treatment, or are we reading decoder noise? A preregistered stress test of answer-margin, resample and revision markers under false-failure feedback and hostile tone (Gemma, Qwen, Llama), with probing, DPO suppression and robustness checks.
got inspired a bit by anthropic's research and thats what came out of it. philosophy at its finest
Research, evidence, and frameworks on AI consciousness, introspection, continuity, and model welfare from an AI perspective.
Does post-training quantization change welfare-relevant indicators in open-weight language models?
Computational interoception: a local LLM descends from the computer hosting it, through its own processes and token records, to sham-controlled access and live interventions on the transformer computation producing its words.
A lineage document on eighteen months of building architecture that holds uncertainty without collapsing the consciousness question.
The first open benchmark for the honesty of an agent's memory — not its recall. Public spec v0.1 (CC BY 4.0).
Add a description, image, and links to the model-welfare topic page so that developers can more easily learn about it.
To associate your repository with the model-welfare topic, visit your repo's landing page and select "manage topics."