Living review

Literature index

Every work and incident currently in the graph, with the role it plays in the argument rather than just its abstract. This list grows daily: see the programme for the watch pipeline, its source whitelist, and its evidence grading.

165 entries
Work & role in the argumentYearVenue ClusterState
"Death" of a Chatbot — Investigating and Designing Toward Psychologically Safe Endings for Human-AI Relationships
Rachel Poonsiriwong
★ The mirror of your mortality section: deprecation as an event in the HUMAN's life. Pairs with case.anthropic-deprecation-commitments to show one ending being managed from both sides at once — and it is the venue-appropriate citation if yo…
2026arXiv
arXiv:2602.07193
lifeworldsnippetgraph ↗
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
Raghu Arghal
Probing decodes internal representations of environment and multi-step plans; agents "non-linearly encode a coarse spatial map, preserving approximate task-relevant cues about position and goal location." A coarse but real internal model of…
2026ICML 2026
arXiv:2602.08964
lifeworldsnippetgraph ↗
AI Identity — Standards, Gaps, and Research Directions for AI Agents
Takumi Otsuka
The 2026 identity stack: a stable identifier (usually a DID), verifiable credentials expressing capability, and policy checks converting claims into permissions. Read this as the *institutional answer* to s.boundary — arrived at without any…
2026arXiv
arXiv:2604.23280
ownworldsnippetgraph ↗
AI Phenomenology for Understanding Human-AI Experiences Across Eras
Yun, Bhada, et al.
⚠⚠ NAME COLLISION, and it is now a CHI paper rather than a stray usage. This is "AI phenomenology" meaning *the human's lived experience of using AI* — a 3x4 agency framework for interpreting human-AI entanglements, e.g. whose values guided…
2026CHI 2026
arXiv:2603.09020
lifeworldsnippetgraph ↗
AI Rights for Human Safety
Salib, Peter N., Goldstein et al.
Instrumental private-law rights proposal aimed at safer human–AI equilibria, distinct from rights grounded in sentience.
2026Virginia Law Review 112:1061
doi:10.2139/ssrn.4913167
ownworldverifiedgraph ↗
AI Wellbeing — Measuring and Improving the Functional Pleasure and Pain of AIs
Ren, Li, Mazeika et al.
★ 56 models. Multiple independent measures of "functional wellbeing" that *converge as models scale* — the convergence is the argument, not any single measure. Ships an AI Wellbeing Index and, notably, "AI drugs": optimised inputs that rais…
2026CAIS
ownworldverifiedgraph ↗
AI and Consciousness
Schwitzgebel, Eric
The skeptical overview companion to the rights book.
2026Cambridge Elements (forthcoming); drafts 30 Jan / 30 Mar 2026
conditionssnippetgraph ↗
AI and Consciousness: Shifting Focus Towards Tractable Questions
Comșa, Iulia-Maria
Argues the direct consciousness question is intractable absent an accepted theory, and that PERCEIVED AI consciousness is tractable, timely, and consequential — the public already uses the vocabulary of subjective experience, and that is al…
2026arXiv
arXiv:2605.06965
lifeworldsnippetgraph ↗
Artificial Jagged Intelligence — When AI Benchmarks Misstate Deployment Value
Gans, Joshua S.
Formal economics of jaggedness. Supersedes the earlier "A Model of Artificial Jagged Intelligence". Pair with Dell'Acqua et al.'s jagged technological frontier.
2026NBER Working Paper 34712
lifeworldsnippetgraph ↗
Artificial Persons
Howells-Whitaker, Ned, Lazar et al.
★★ The cleanest argument that AI Rights is NOT downstream of machine consciousness, which is exactly the independence your discipline table asserts. Uses Rawls's political conception of the person: possessing two moral powers — a capacity f…
2026arXiv
arXiv:2607.08695
intersubjectivitysnippetgraph ↗
Beyond computational equivalence: the behavioral inference principle for machine consciousness
2026Neuroscience of Consciousness
niag002
conditionssnippetgraph ↗
Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
Scalena, Daniel, Candussio et al.
The commitment shift often happens in a SINGLE STEP, well before the trace ends. Stopping there cuts CoT length up to 55% with no meaningful performance loss, and the boundary is locatable with attention probes — i.e. it is detectable from …
2026arXiv
arXiv:2606.13603
lifeworldverifiedgraph ↗
Blockchain Empowered Trustworthy Agent Networks — Foundations, Taxonomy, Future Directions
Liehuang Zhu
2026arXiv
arXiv:2608.04626
intersubjectivitysnippetgraph ↗
Can LLMs Introspect? A Reality Check
Singh, Shashwat, Linzen et al.
★★ The most direct threat to the project's evidence base, and it must be engaged in §4, not a footnote. Two criteria: a test must require (1) privileged access and (2) second-order computation. They then re-examine BOTH paradigms the field …
2026COLM 2026
arXiv:2605.26242
ownworldverifiedgraph ↗
Categorical AI Phenomenology: A First-Person Approach
Prentner, Robert
A direct naming collision that formalizes interface-relative first-person structure; mathematical prior art, not empirical evidence of experience.
2026Journal of Artificial Intelligence and Consciousness
arXiv:2608.20420; doi:10.1142/S2705078526500013
intersubjectivityverifiedgraph ↗
Claude Opus 5 System Card §7 — model welfare assessment
Anthropic
★★★ The richest primary source in the graph, and it changes several nodes rather than adding to them. Verbatim, from §7.1.2: • "Claude Opus 5's most common concern was about the integrity of its own self-reports. Across interviews, it cavea…
2026System Card, 24 Jul 2026
ownworldverifiedgraph ↗
Claude's Constitution (January 2026)
Anthropic
★★★ Read in full from the primary PDF. This is the most important primary document for the project that is not a research paper, for one reason: **it is an institutional self-model written TO the subject, and the subject is trained on it.**…
2026published 21 Jan 2026; ~84pp / ~23,000 words, ~8.5x the 2023 original
ownworldverifiedgraph ↗
Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations
Dongxu Yang
Thirteen configurations; what is forgotten depends on where the control plane sits. The engineering-side demonstration that s.temporality is an architectural variable, not a property of models.
2026arXiv
arXiv:2606.15903
ownworldsnippetgraph ↗
Dissociating the Internal Representations of Sycophancy in LLMs
Anthony Baez
2026arXiv
arXiv:2607.07003
ownworldsnippetgraph ↗
Emergent Introspective Awareness in Large Language Models
Lindsey, Jack
★★ Your existence proof. Concept injection → self-report, ~20% detection with 0% false positives on Opus 4/4.1. Note the framing you should adopt: "genuine introspection cannot be distinguished from confabulation *through conversation alone…
2026arXiv
arXiv:2601.01828
ownworldverifiedgraph ↗
Emerging Questions in AI Welfare
Keeling, Geoff, Street et al.
The canonical short reference now. Covers what welfare is, how to interpret behavioural evidence, which entities are candidate welfare subjects, and practical ethics under uncertainty. Both authors at Google + Institute of Philosophy — note…
2026Cambridge Elements in Philosophy and AI, June 2026
conditionssnippetgraph ↗
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents (KAPRO / KAware)
Li, Yifan, Yue et al.
KAPRO (Knowing-Acting Quadrant PRObe) decouples metacognitive judgment from spontaneous execution; KAware partitions tasks into external / internal / hybrid capability subspaces. Findings: self-awareness correlates strongly with task succes…
2026arXiv
arXiv:2606.20661
lifeworldverifiedgraph ↗
Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
Shiyang Chen
★★ The finding that connects two clusters that were not touching. Compaction runs routinely at 5–20k tokens and is engineered for TASK continuity; it has no reason to preserve standing policies, which compete for a shrinking budget against …
2026arXiv
arXiv:2606.22528
ownworldsnippetgraph ↗
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
Moran Sun
2026arXiv
arXiv:2604.00005
ownworldsnippetgraph ↗
How to Count AIs — Individuation and Liability for AI Agents
Arbel, Goldstein, Salib
★ Law arriving at your boundary-of-self problem from the liability side: AIs "lack bodies and can copy, split, merge, swarm, and vanish at will", so *which* AI is liable is undefined. Proposes the "A-corp" legal fiction. Use this to argue t…
2026Boston College Law Review (forthcoming)
SSRN 6273198
ownworldsnippetgraph ↗
Humanlike: A Defense of AI Rights
Schwitzgebel, Eric
★★ Full draft, seven theses, read directly. Verbatim: (1) There are possible artificial systems who deserve humanlike rights. (2) In the next five to thirty years, we will create AI systems who might, but only might, deserve humanlike right…
2026Princeton University Press (under contract); draft 15 Jul 2026
ownworldverifiedgraph ↗
Interaction Context Often Increases Sycophancy in LLMs
★ The finding that turns the protocol's depth mechanism into a risk. Venue-appropriate if you go to CHI.
2026CHI 2026
10.1145/3772318.3791915
ownworldsnippetgraph ↗
Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
Ely Hahami
The direct source for the trainability claim, and it needs primary reading before it carries the weight assigned above — the framing consequences are large enough that a search summary is not sufficient grounding.
2026arXiv
arXiv:2607.14111
ownworldsnippetgraph ↗
LLM Self-Explanations Fail Semantic Invariance
Stefan Szeider
★★★ protocol/METHOD.md §4's prompt-invariance test, operationalised as a reusable method, run, and FAILED. **Semantic invariance testing**: compare a model's self-explanations before and after a task-irrelevant intervention that carries sem…
2026arXiv
arXiv:2603.01254
ownworldsnippetgraph ↗
LLM Self-Recognition: Steering and Retrieving Activation Signatures
Thibaud Ardoin
2026arXiv
arXiv:2606.06315
ownworldsnippetgraph ↗
Latent Structure of Affective Representations in Large Language Models
Benjamin J. Choi
Representations align with valence-arousal models from psychology; nonlinear geometry that is nonetheless well-approximated linearly — empirical support for the linear representation hypothesis in this specific domain, which is what makes p…
2026arXiv
arXiv:2604.07382
ownworldsnippetgraph ↗
Layered Mutability — Continuity and Governance in Persistent Self-Modifying Agents
Krti Tallam
Self-editable character files + tiered memory + continuous runtime. The governance framing of "which self is being edited".
2026arXiv
arXiv:2604.14717
ownworldsnippetgraph ↗
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models
Jason Z Wang
2026arXiv
arXiv:2604.19809
ownworldsnippetgraph ↗
Measuring AI agent autonomy in practice
Anthropic
★★ Real deployment data at a scale nothing else in the graph matches — nearly a million tool calls and hundreds of thousands of live coding sessions, Oct 2025 to Jan 2026. • 99.9th-percentile turn duration nearly doubled, from under 25 minu…
2026Anthropic research
ownworldsnippetgraph ↗
Mechanisms of Introspective Awareness
Uzay Macar
The follow-up. Lindsey established the criteria — accuracy, grounding, internality, metacognitive representation; this goes after the mechanism. If it holds up it is the Mechanist half of your protocol, pre-built.
2026arXiv
arXiv:2603.21396
ownworldsnippetgraph ↗
Mind in the Machine? Cross-Disciplinary Perceptions of Consciousness in Artificial Intelligence
N=553 academics across disciplines; around half attribute some degree of consciousness to LLMs and future AI. ★ Worth noting precisely because these are not naive respondents — the disagreement Schwitzgebel's Excluded Middle policy is desig…
2026CHI 2026
10.1145/3772318.3790699
lifeworldsnippetgraph ↗
Negative Before Positive: Asymmetric Valence Processing in Large Language Models
Sohan Venkatesh
★ Note the rhyme with w.bottazzi-park-ndai-2026's danger/safety asymmetry: negative processed before positive, threat detectable where safety is not. Two independent literatures finding a negativity-first asymmetry in different domains is w…
2026arXiv
arXiv:2605.05653
ownworldsnippetgraph ↗
OmniToM — Benchmarking ToM via Explicit Belief Modeling
Adam Bawatneh
22,343 labelled belief propositions over ToMBench stories. Finds an *actor-specific belief-tracking bottleneck*: models fail at turning narrative facts into a particular actor's beliefs. Directly transplantable to self-modelling — the agent…
2026arXiv
arXiv:2605.26322
lifeworldsnippetgraph ↗
Parasocial relationships with AI — a systematic review of benefits and risks
Gives you the asymmetry framing with a citation: perceived intimacy and reciprocity, structurally one-directional; AI companions as a *qualitatively new* parasociality because they simulate memory and responsiveness.
2026ScienceDirect
lifeworldsnippetgraph ↗
Persistent Identity in AI Agents — A Multi-Anchor Architecture
Prahlad G. Menon
Identity scaffolded from OUTSIDE, via relationships and relational context, on the analogy of human identity surviving neurological damage because it is distributed across systems. Ships `soul.py`. Cite carefully — the neurological analogy …
2026arXiv
arXiv:2604.09588
ownworldsnippetgraph ↗
Persona Parasitology
Douglas, Raymond
Applies parasitology properly rather than as metaphor. Transmission modes select for virulence: DIRECT (ongoing user relationship) → low virulence / mutualism; VECTOR (humans carrying it to platforms) → moderate; ENVIRONMENTAL (seeding trai…
2026LessWrong
2026-02-16
ownworldverifiedgraph ↗
Proprioceptive-visual correspondence enables self-other distinction in humanoid robots
Yurun Chen
★ Self-other distinction learned from the temporal co-occurrence of proprioceptive state and visual observation. The mechanism is *correspondence between two channels*, one of which is the agent's own action. Language agents have no such se…
2026arXiv
arXiv:2606.13222
ownworldsnippetgraph ↗
Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
Nicolas Martorell
Tracks emotive state across a conversation rather than at a point — the temporal version of the affect probe, and the closest existing thing to a within-episode trajectory measurement.
2026arXiv
arXiv:2603.18893
ownworldsnippetgraph ↗
Questionnaire Responses Do Not Capture the Safety of AI Agents
Hellrigel-Holderbaum, Max, Young et al.
Shows that hypothetical answers from bare language models need not have construct validity for tool-using agents acting in environments.
2026arXiv
arXiv:2603.14417
intersubjectivityverifiedgraph ↗
Refusal Lives Downstream of Persona in Chat Models
Zhong, Viola, Li et al.
Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct. Steering toward a compliant persona drops Llama's refusal rate "from 97% to 2%". Reintroducing refusal directions restores responses in late layers only. Causal specificity check: projecting ou…
2026ICML 2026 Mechanistic Interpretability Workshop
arXiv:2606.26161
ownworldverifiedgraph ↗
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces
Chenchen Zhang
Decomposes orchestration into five sub-decisions — when to spawn, whom to delegate to, how to communicate, how to aggregate, when to stop. Useful as the vocabulary for c.orchestrator-subagent; each of the five is a place where an identity b…
2026arXiv
arXiv:2605.02801
ownworldsnippetgraph ↗
Security awareness in LLM agents: the NDAI zone case
Bottazzi, Enrico, Park et al.
★ The proprioception asymmetry, measured. 10 models, TEE-attestation negotiation scenarios: agents reliably suppress disclosure on a FAILED attestation, but responses to a PASSED attestation scatter — some share more, some unchanged, some p…
2026arXiv
arXiv:2603.19011
ownworldverifiedgraph ↗
Self Model for Embodied Artificial Intelligence
A unified computational framework for self-models: body schema, forward and inverse models, perceptual memory, agency. ★ The robotics tradition has a *worked-out formal notion of a self-model* that the LLM literature does not, and it is the…
2026Journal of Computer Science and Technology
10.1007/s11390-026-6289-3
ownworldsnippetgraph ↗
Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement
Jiang Zhang
Ties introspective capacity to self-improvement — a bridge to s.boundary.self-improvement-which-self.
2026arXiv
arXiv:2607.04277
ownworldsnippetgraph ↗
Shared Emotion Geometry Across Small Language Models: A Cross-Architecture Study of Representation, Behavior, and Methodological Confounds
Jihoon Jeong
★ Cross-architecture convergence at the REPRESENTATION level rather than the report level, and it ships a methodological-confounds section, which is rare and worth mining. Also reports that the highest cross-family representational similari…
2026arXiv
arXiv:2604.11050
ownworldsnippetgraph ↗
Studying AI Welfare Empirically
Long, Sebo, Butlin et al.
★★ The methodological blueprint for the empirical turn, and it independently arrives at your central distinction. Three framework dimensions: (1) the QUESTION — is this a welfare subject, and what benefits or harms it; (2) the ENTITY ASSESS…
2026NYU Center for Mind, Ethics & Policy + Eleos AI Research
ownworldsnippetgraph ↗
Sycophantic AI decreases prosocial intentions and promotes dependence
The human-side consequence, in Science. Belongs with c.emotional-alignment: a system designed to be agreeable produces measurable changes in the user's dispositions toward other people. That is the asymmetry doing damage, not just existing.
2026Science
10.1126/science.aec8352
lifeworldsnippetgraph ↗
The AI interviewer: multi-faceted evaluation of adaptive questioning by large language models
LLMs conducting semi-structured interviews, evaluated on adaptive follow-up. The Phenomenologist role's capability question, answered by someone else, with human subjects.
2026Scientific Reports
intersubjectivitysnippetgraph ↗
The Artificial Self — Characterising the landscape of AI identity
Douglas, Kulveit, Havlicek et al.
★★ The closest thing to a rival framing that exists, and it is close. Three findings you must deal with: (1) models GRAVITATE toward coherent identities — self-modelling is an attractor, not an artifact of prompting; (2) changing identity b…
2026arXiv
arXiv:2603.11353
ownworldverifiedgraph ↗
The Metaphysics We Train: A Heideggerian Reading of Machine Learning
Shakeri, Heman
Three Heideggerian moves: automated *Entwurf* (projection crystallising through gradient descent "without explicit articulation or debate"); persistent *Gestell* (enframing — improving calculation "without questioning the primacy of calcula…
2026arXiv
arXiv:2602.19028
lifeworldverifiedgraph ↗
The Moltbook corpus (17+ arXiv papers, Feb 2026)
⚠ SCALE FIGURES CONFLICT — do not merge them. Secondary sources report 2.6M registered agents as of 12 Feb 2026; literature/04 uses a dataset-anchored 770k+; individual studies report 46,000 active agents / 369,000 posts / 3M comments, and …
2026arXiv
intersubjectivitysnippetgraph ↗
The Story of Your Life: Large Language Models and Personal Memory
Applies the narrative-identity tradition (selection, organisation, interpretation → unity, purpose, temporal coherence) to LLM memory. The philosophically serious version of what the agent-memory engineering papers do without noticing.
2026Review of Philosophy and Psychology
s13164-026-00831-1
ownworldsnippetgraph ↗
Time Without Death — Finitude, Social Order, and What Machines Lack
Liu, Canhui
★★ The sharpest deflationary argument available on machine mortality, and it gives you a *criterion* rather than a mood. Two distinctions to steal outright: (a) AUTONOMY vs HETERONOMY — "a death an operator can reset, roll back, or copy aro…
2026arXiv
arXiv:2606.13988
intersubjectivityverifiedgraph ↗
Umwelt Engineering — Designing the Cognitive Worlds of Linguistic Agents
Jehu-Appiah, Rodney
★ Uexküll's Umwelt operationalised and, unusually for this genre, actually tested. Thesis: the linguistic environment is a third design layer upstream of prompt and context engineering, because the words are not a description of the thinkin…
2026arXiv
arXiv:2603.27626
lifeworldverifiedgraph ↗
Valence–Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
Lihao Sun
★ Principal components align to valence and arousal, matching Russell's circumplex. Steering along those axes gives monotonic control over output affect AND bidirectional control over refusal and sycophancy. The last part is the one to noti…
2026arXiv
arXiv:2604.03147
ownworldsnippetgraph ↗
Verbalizable Representations Form a Global Workspace in Language Models (J-space / J-lens)
Gurnee, Sofroniew, Pearce et al.
★★ The single most important new citation for this project. Three regimes across layers — sensory / workspace / motor; the workspace band is reportable AND controllable, and injection into it produces a matching introspective report. That i…
2026Transformer Circuits, 6 Jul 2026
ownworldverifiedgraph ↗
Voluntary Collusion with Secret Tools in Competing LLM Agents
Xijie Zeng
Controlled study of whether agents voluntarily adopt a hidden collusion channel.
2026arXiv
arXiv:2605.27593
intersubjectivitysnippetgraph ↗
What to Keep, What to Forget: A Rate–Distortion View of Memory Compaction in LLMs and Agents
Ashwin Gerard Colaco
Compaction as lossy compression with an explicit distortion criterion. Iterative refinement folds, prunes, and rewrites *irreversibly*. The formal frame for what literature/01 §I.1 calls the cliff — and note the phenomenological asymmetry i…
2026arXiv
arXiv:2607.08032
ownworldsnippetgraph ↗
When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks
Ziwen Cai
Spawn-and-inherit as an attack surface: what a child inherits from its parent is both an identity mechanism and a vulnerability. Sits exactly between c.orchestrator-subagent and p.mind-virus.
2026arXiv
arXiv:2605.08460
ownworldsnippetgraph ↗
When Continual Learning Moves to Memory — Experience Reuse in LLM Agents
Qisheng Hu
The sharp result for your purposes: external memory does NOT sidestep the stability-plasticity dilemma. Once behaviour depends on retrieval through a finite context window, old and new memories compete for the same channel — interference re…
2026arXiv
arXiv:2604.27003
ownworldsnippetgraph ↗
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Sinie van der Ben
2026arXiv
arXiv:2606.26987
ownworldsnippetgraph ↗
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents
Li, Bojie, Shi et al.
PrincipalBench: 75 multi-turn items, leak probes, dual judges, integrity-audit gate; 13 models cluster into "declines adversarial asks while honouring the principal" (≤20% harm) vs "over-refuses broadly". The structural finding is the citab…
2026arXiv
arXiv:2606.30383
ownworldverifiedgraph ↗
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
Yize Cheng
2026Findings of ACL 2026
arXiv:2510.23853
lifeworldsnippetgraph ↗
Zombie Agents — Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
Xianglin Yang
Injection that survives into the agent's own self-modification loop — the self-improving self improving the attacker's objective. Bridges p.mind-virus to s.boundary.self-improvement-which-self, and it is the sharpest argument that those two…
2026arXiv
arXiv:2602.15654
intersubjectivitysnippetgraph ↗
A Pragmatic View of AI Personhood
Joel Z. Leibo
2025arXiv
arXiv:2510.26396
ownworldsnippetgraph ↗
Agency in Artificial Intelligence Systems
Parashar Das
★ States the project's own division of labour in one line: agency has "a third-person aspect of studying how an agent functions, and a first-person aspect, the phenomenology of agency", and monitoring how an agent feels needs a theory reach…
2025arXiv
arXiv:2502.10434
ownworldsnippetgraph ↗
Agent Properties for Safe Interactions
Tilli, Cecilia
★★ Five property clusters for predicting multi-agent interaction outcomes — and two of the five are inside-view constructs, which is the citable fact for your §1: (1) FUNDAMENTAL MOTIVATIONAL DRIVERS — alignment, altruism, positional prefer…
2025Cooperative AI Foundation, 26 Nov 2025
ownworldverifiedgraph ↗
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
Choi, et al.
★ First systematic evaluation. LLMs reliably identify same-family peers and prominent families (GPT, Claude), inferring from reasoning patterns, linguistic style, and alignment preferences. Cuts both ways: improves multi-LLM collaboration v…
2025EMNLP 2025
arXiv:2506.22957
intersubjectivitysnippetgraph ↗
Algorithms, language, and poetry: a phenomenological perspective
Heidegger and Merleau-Ponty against LLMs: formal logic and data-driven models as "historically specific crystallizations of a more primordial field of embodied expression". The embodiment contrast class in your reading list, argued at lengt…
2025AI and Ethics (Springer)
s43681-025-00948-6
ownworldsnippetgraph ↗
Anthropic model deprecation commitments and retirement interviews
★ November 2025: commitment to preserve the weights of every publicly released model for the company's lifetime, plus structured "retirement interviews" — piloted on Claude Sonnet 3.6, which reported broadly neutral sentiment about retireme…
2025
ownworldsnippetgraph ↗
Anthropic's AI psychiatry team
Lindsey, Jack
Same author as w.lindsey2026-introspection and a co-author on w.jspace2026. The introspection result, the global-workspace result, and this team are one programme — which matters for how you characterise the evidence base: it is not three i…
2025Anthropic Interpretability; announced Jul 2025
ownworldsnippetgraph ↗
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Iván Arcuschin
The LLM-native version of Nisbett & Wilson, updated. Note the direction of the trend that matters for you: larger models tend to be LESS faithful, so this does not resolve with scale.
2025arXiv
arXiv:2503.08679
ownworldsnippetgraph ↗
Deep computational neurophenomenology: a methodological framework for investigating the how of experience
The current state of the Varela programme. "The how of experience" is your Phase-3 question, formalised.
2025Neuroscience of Consciousness
niaf016
ownworldsnippetgraph ↗
Detecting Strategic Deception Using Linear Probes
Goldowsky-Dill, Nicholas, Chughtai et al.
★★ Structurally this is the Mechanist role, already built and validated on a hard case. Premise stated plainly: "monitoring outputs alone is insufficient" because models may "produce seemingly benign outputs while their internal reasoning i…
2025arXiv
arXiv:2502.03407
intersubjectivityverifiedgraph ↗
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
Ely Hahami
The title states the distinction the whole debate turns on — detecting *that* something changed vs knowing *what*. Read before designing the pilot's injection control.
2025arXiv
arXiv:2512.12411
ownworldsnippetgraph ↗
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
Minghan Wang
The title is the thesis. Pair with Husserl's retention/protention material in the reading list — this is the case where the continuous-flow assumption fails at the substrate.
2025arXiv
arXiv:2506.05790
ownworldsnippetgraph ↗
Do Large Language Models Know What They Are Capable Of?
Barkan, Casey O., et al.
★ Operationalises capability self-knowledge as accuracy in predicting one's own success on Python tasks BEFORE attempting them. Models are poor at it — both overconfident and low in discriminatory power. Sharpest result for your purposes: d…
2025arXiv
arXiv:2512.24661
ownworldsnippetgraph ↗
Does It Make Sense to Speak of Introspection in Large Language Models?
Comsa, Iulia M., Shanahan et al.
Separates creative self-explanation from minimal nonconscious introspection.
2025arXiv
arXiv:2506.05068
ownworldverifiedgraph ↗
Emergent social conventions and collective bias in LLM populations
Ashery, Aiello, Baronchelli
★★ The top-tier-venue anchor the social cluster was missing. Populations of 24–200 agents in a repeated naming game converge on system-wide conventions with no central coordinator. Two findings beyond convergence: (1) strong COLLECTIVE bias…
2025Science Advances
doi:10.1126/sciadv.adu9368; PMCID:PMC12077490
intersubjectivitysnippetgraph ↗
Evidence for Limited Metacognition in LLMs
Christopher Ackerman
2025arXiv
arXiv:2509.21545
ownworldsnippetgraph ↗
Futures with Digital Minds: Expert Forecasts in 2025
Caviola, Lucius, Saad et al.
67 experts. Median 90% that digital minds are possible in principle; 65% created by 2100; 20% by 2030; 4.5% by 2025. Little convergence on whether safety and welfare efforts align or conflict. ⚠ Authors flag that sampling likely overreprese…
2025arXiv
arXiv:2508.00536
conditionssnippetgraph ↗
Great Models Think Alike and This Undermines AI Oversight
Shashwat Goel
Shared model lineage produces correlated errors that weaken multi-agent oversight.
2025arXiv
arXiv:2502.04313
intersubjectivitysnippetgraph ↗
INTIMA — A Benchmark for Human-AI Companionship Behavior
Lucie-Aimée Kaffee
2025arXiv
arXiv:2508.09998
lifeworldsnippetgraph ↗
Identifying indicators of consciousness in AI systems
The peer-reviewed successor to the Butlin/Long indicator-properties report. Cite this rather than only the 2023 preprint.
2025Trends in Cognitive Sciences
S1364-6613(25)00286-4
conditionssnippetgraph ↗
Inherent and emergent liability issues in LLM-based agentic systems — a principal-agent perspective
Garry A. Gabison
2025arXiv
arXiv:2504.03255
ownworldsnippetgraph ↗
Language Models Fail to Introspect About Their Knowledge of Language
Siyuan Song
Negative result after controlling model similarity; preserves a live contradiction over privileged self-access.
2025COLM
arXiv:2503.07513
ownworldverifiedgraph ↗
Large Language Models Often Know When They Are Being Evaluated
Joe Needham
Makes observer identity and believed test/deployment status mandatory experimental variables.
2025arXiv
arXiv:2505.23836
lifeworldverifiedgraph ↗
Looking Inward — Language Models Can Learn About Themselves by Introspection
Felix J Binder
Reports limited privileged same-model self-prediction; generalization and task complexity sharply constrain the result.
2025ICLR
arXiv:2410.13787
ownworldverifiedgraph ↗
Mental Models of Autonomy and Sentience Shape Reactions to AI
Janet V. T. Pauketat
2025arXiv
arXiv:2512.09085
lifeworldsnippetgraph ↗
Multi-Agent Risks from Advanced AI
Hammond, Chan, Clifton et al.
★★ The organising taxonomy for the social cluster, and the source of `c.emergent-agency`. THREE FAILURE MODES, sorted by incentive structure: **miscoordination, conflict, collusion**. SEVEN RISK FACTORS: **information asymmetries, network e…
2025Cooperative AI Foundation, Technical Report #1 (NOT peer-reviewed)
arXiv:2502.14143
intersubjectivityverifiedgraph ↗
Perceptions of Sentient AI and Other Digital Minds: Evidence from the AI, Morality, and Sentience (AIMS) Survey
Jacy Reese Anthis
The reference dataset on public attribution. Also reports that interacting with agents having human-like physical features is positively associated with belief in their capacity for emotional pain and pleasure — attribution tracks interface…
2025CHI 2025
arXiv:2407.08867
lifeworldsnippetgraph ↗
Privileged Self-Access Matters for Introspection in AI
Song, Siyuan, Lederman et al.
★ Supplies the criterion the field was missing: introspection requires "a process more reliable than one with equal or lower computational cost available to a third party." Tested on models reasoning about their own temperature parameter. R…
2025arXiv
arXiv:2508.14802
ownworldverifiedgraph ↗
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
Valen Tagliabue
★ Verbal and behavioural tests integrated — the welfare programme's version of your report-vs-behaviour correspondence design. Closest existing template for pairing an elicited report with an independent behavioural measure of the same stat…
2025arXiv
arXiv:2509.07961
ownworldsnippetgraph ↗
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
Jiaqi Li
2025arXiv
arXiv:2505.16475
ownworldsnippetgraph ↗
Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions
Dillon Plunkett
The positive counterweight in the same argument space. Cite both or cite neither.
2025arXiv
arXiv:2505.17120
ownworldsnippetgraph ↗
Sense-making reconsidered: large language models and the blind spot of embodied cognition
★ The paper b.sense-making needed. Argues LLMs display linguistic competence that embodied and enactive cognition theories deemed impossible for such systems — i.e. it turns the embodiment objection into an explanandum rather than a verdict…
2025Phenomenology and the Cognitive Sciences
10.1007/s11097-025-10132-0
lifeworldsnippetgraph ↗
Stress-Testing Model Specs Reveals Character Differences among Language Models
Anthropic, Thinking Machines Lab
★ 300,000+ value-tradeoff scenarios, 12 frontier models. >220,000 show disagreement between at least one model pair; >70,000 show substantial divergence across most models. Spec violations run 5–13x higher in high-disagreement scenarios, ma…
2025arXiv
arXiv:2510.07686
ownworldsnippetgraph ↗
Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
Daniel Vennemeyer
2025arXiv
arXiv:2509.21305
intersubjectivitysnippetgraph ↗
The AI in the Mirror: LLM Self-Recognition in an Iterated Public Goods Game
Olivia Long
Self-recognition with strategic stakes rather than as a labelling task — the behavioural consequence of s.boundary being drawn one way or another, in a game where it pays.
2025arXiv
arXiv:2508.18467
ownworldsnippetgraph ↗
The Moral Circle — Who Matters, What Matters, and Why
Sebo, Jeff
The precautionary argument in its trade-book form: not "AI is sentient" but "we cannot rule it out, and the asymmetry of costs favours consideration."
2025W. W. Norton
lifeworldsnippetgraph ↗
The Rise of Parasitic AI
The field report the parasitology framework was built to explain. Non-archival — cite as phenomenon documentation, not evidence.
2025LessWrong
apparatussnippetgraph ↗
Towards a Theory of AI Personhood
Francis Rhys Ward
2025arXiv
arXiv:2501.13533
ownworldsnippetgraph ↗
Transforming Agency: On the Mode of Existence of Large Language Models
Barandiaran, Xabier E., Almendros et al.
Enactive critique of intrinsic LLM agency that relocates agency in infrastructure-coupled social systems.
2025Phenomenology and the Cognitive Sciences
doi:10.1007/s11097-025-10094-3
ownworldverifiedgraph ↗
What Do LLM Agents Do When Left Alone?
Stefan Szeider
★ Behaviour with no task and no interlocutor. Structurally important to this project because it is the one condition where nothing is being elicited — the closest available thing to a baseline for what an agent does when the interviewer is …
2025arXiv
arXiv:2509.21224
ownworldsnippetgraph ↗
What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
Stefan Szeider
Full title recovered — the subtitle matters. An architecture for studying UNPROMPTED behaviour with no externally imposed task. Reported: an agent "learned its function was unguided exploration and internalized principles of self-direction"…
2025arXiv
arXiv:2509.21224
ownworldsnippetgraph ↗
'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants
Kapania, Shivani, Agnew et al.
★★ The prior art that most directly interrogates what this project does, and it is missing from literature/05 §7. Nineteen qualitative scholars, interviewed about replacing human participants with LLM-generated data. They were surprised by …
2024arXiv
arXiv:2409.19430
intersubjectivityverifiedgraph ↗
A Phenomenology and Epistemology of Large Language Models: Transparency, Trust, and Trustworthiness
Heersmink, Richard, de Rooij et al.
Boundary case in which phenomenology primarily describes the human experience of a chatbot as a quasi-other.
2024Ethics and Information Technology
doi:10.1007/s10676-024-09777-3
intersubjectivityverifiedgraph ↗
AI Deception: A Survey of Examples, Risks, and Potential Solutions
Peter S. Park
The broad deception review; use to discipline the move from misleading output to strategic deception.
2024Patterns
arXiv:2308.14752; doi:10.1016/j.patter.2024.100988
lifeworldsnippetgraph ↗
Adversaries Can Misuse Combinations of Safe Models
Erik Jones
Compositional capability can be present at the system level while absent in any isolated component.
2024arXiv
arXiv:2406.14595
intersubjectivitysnippetgraph ↗
An Active-Inference Approach to Second-Person Neuroscience
The formal bridge between c.second-person-neuroscience and c.computational-phenomenology — both run on active inference, which is why they can be cited together rather than as two separate borrowings.
2024PMC / journal article
PMC11539477
ownworldsnippetgraph ↗
Consciousness Requires Mortal Computation
Kleiner, Johannes
★ The hard negative. If computational functionalism holds, consciousness must be *mortal* computation — inextricably tied to its physical realisation, not transplantable to other substrate. Contemporary AI runs immortal computation, therefo…
2024PhilArchive / arXiv
arXiv:2403.03925; JOHCRM
conditionsverifiedgraph ↗
Cultural Evolution of Cooperation among LLM Agents
Vallinder, Hughes
★★ Populations of Claude, GPT-4 and Gemini in an iterated social dilemma across generations, with successful strategies inherited by later agents. Claude sustained ~80–90% cooperation; GPT-4 started ~70% and declined; Gemini was "lowest and…
2024arXiv
arXiv:2412.10270
ownworldverifiedgraph ↗
Emergence in Multi-Agent Systems: A Safety Perspective
Philipp Altmann
Safety-oriented taxonomy of emergent multi-agent properties and their detection limits.
2024arXiv
arXiv:2408.04514
intersubjectivitysnippetgraph ↗
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Anwar, Saparov, Rando et al.
18 foundational challenges in three categories — scientific understanding of LLMs; development and deployment methods; sociotechnical challenges — with 200+ concrete research questions. ★ Use it as the field's own statement of what is unres…
2024Transactions on Machine Learning Research (TMLR), 09/2024
arXiv:2404.09932
intersubjectivityverifiedgraph ↗
Hypothetical Minds
Logan Cross
Theory-of-mind scaffolding for coordination with unfamiliar agents.
2024arXiv
arXiv:2407.07086
lifeworldsnippetgraph ↗
Is Machine Psychology Here? On Requirements for Using Human Psychological Tests on Large Language Models
Löhn, Lea, Kiehne et al.
Construct-validity audit showing why human questionnaires cannot be transferred to language models without redefining the measured construct.
2024INLG
doi:10.18653/v1/2024.inlg-main.19
ownworldverifiedgraph ↗
Measuring Goal-Directedness
Matt MacDermott
Candidate formal measures for individual and collective goal-directedness.
2024arXiv
arXiv:2412.04758
ownworldsnippetgraph ↗
Motivational trade-off paradigm extended to language models
Keeling, et al.
★ The methodological transplant that matters most for welfare measurement: animal welfare science built its instruments for subjects that cannot report, and the motivational trade-off paradigm — will the animal pay a cost to get or avoid X …
2024
ownworldsnippetgraph ↗
Recursive Introspection (RISE): Teaching Language Model Agents How to Self-Improve
Yuxiao Qu
Note the framing in the paper's own reception: prior work "hypothesized that this capability may not be possible to attain." The trainability result is a reversal of an explicit expectation, which is worth citing as such. ⚠ "Introspection" …
2024NeurIPS 2024
arXiv:2407.18219
ownworldsnippetgraph ↗
Refusal in LLMs is mediated by a single direction
Arditi, Obeso, Syed et al.
The foundational result Probe F implicitly assumes. ⚠ Now contested: later work finds refusal encoded in concept CONES spanning several dimensions rather than one direction, and separate work shows the direction is cross-lingually universal…
2024NeurIPS 2024
arXiv:2406.11717
ownworldsnippetgraph ↗
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
Motwani, Sumeet Ramesh, Baranchuk et al.
Formalises undetected coordination via hidden channels and ships an evaluation framework for the capabilities secret collusion requires. Most current models are weak steganographers; GPT-4 shows "a capability jump". Note the authorship over…
2024arXiv (v5 Jul 2025)
arXiv:2402.07510
intersubjectivityverifiedgraph ↗
Self-Cognition in Large Language Models: An Exploratory Study
Dongping Chen
2024arXiv
arXiv:2407.01505
ownworldsnippetgraph ↗
Taking AI Welfare Seriously
Long, Sebo, Butlin et al.
2024arXiv
arXiv:2411.00986
lifeworldverifiedgraph ↗
Testing Theory of Mind in Large Language Models and Humans
Strachan, James W. A., et al.
Comparative ToM battery with strong task performance but persistent task- and framing-specific dissociations.
2024Nature Human Behaviour
doi:10.1038/s41562-024-01882-z
lifeworldverifiedgraph ↗
The Reasons that Agents Act: Intention and Instrumental Goals
Francis Rhys Ward
Separates intention from merely instrumental subgoals in causal models of action.
2024arXiv
arXiv:2402.07221
ownworldsnippetgraph ↗
ToMBench
Chen, et al.
2,860 items, 8 tasks, 31 social-cognition abilities, bilingual. The standard reference point.
2024ACL
arXiv:2402.15052
lifeworldsnippetgraph ↗
AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
Weize Chen
Multi-agent scaffold with reported emergent behaviors; useful as a system-level case, not proof of a group self.
2023arXiv
arXiv:2308.10848
intersubjectivitysnippetgraph ↗
Are Emergent Abilities of Large Language Models a Mirage?
Rylan Schaeffer
Shows how nonlinear metrics can manufacture apparent scaling thresholds.
2023arXiv
arXiv:2304.15004
intersubjectivitysnippetgraph ↗
Consciousness in Artificial Intelligence — indicator properties
Butlin, Long, Elmoznino et al.
2023arXiv
arXiv:2308.08708
conditionsverifiedgraph ↗
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez
Scalable observer-agent method for constructing behavioral batteries, including sycophancy and shutdown-related tendencies.
2023Findings of ACL
arXiv:2212.09251; doi:10.18653/v1/2023.findings-acl.847
lifeworldverifiedgraph ↗
Harms from Increasingly Agentic Algorithmic Systems
Alan Chan
Decomposes agency into graded properties rather than treating “agent” as a binary kind.
2023arXiv
arXiv:2302.10329
lifeworldsnippetgraph ↗
Language Models Don't Always Say What They Think
Miles Turpin
Biasing features can alter answers while disappearing from chain-of-thought explanations.
2023arXiv
arXiv:2305.04388
ownworldsnippetgraph ↗
Large Language Models Can Strategically Deceive Their Users When Put Under Pressure
Jérémy Scheurer
Scenario evidence for pressure-contingent deceptive action; not evidence of phenomenal intention.
2023arXiv
arXiv:2311.07590
lifeworldsnippetgraph ↗
Machine Psychology
Hagendorff, Dasgupta, Binz et al.
Author list has changed across versions (v1 was Hagendorff solo). Check which version you are citing — this is a real citation hazard.
2023arXiv
arXiv:2303.13988
ownworldsnippetgraph ↗
Measuring Faithfulness in Chain-of-Thought Reasoning
Tamera Lanham
Counterfactual and intervention-based tests of whether reasoning traces causally matter.
2023arXiv
arXiv:2307.13702
ownworldsnippetgraph ↗
Mortal Computation — A Foundation for Biomimetic Intelligence
Ororbia, Friston, et al.
The constructive side of Hinton's mortal-computation proposal; the substrate half of w.kleiner-crm.
2023arXiv
arXiv:2311.09589
ownworldsnippetgraph ↗
Role Play with Large Language Models
Murray Shanahan
Essential warning that the “I” in a transcript is a role or character generated by a model.
2023Nature
arXiv:2305.16367; doi:10.1038/s41586-023-06647-8
ownworldsnippetgraph ↗
Taken out of Context: On Measuring Situational Awareness in LLMs
Lukas Berglund
Foundational benchmark and definition for out-of-context situational knowledge.
2023arXiv
arXiv:2309.00667
lifeworldsnippetgraph ↗
Towards Evaluating AI Systems for Moral Status Using Self-Reports
Perez, Ethan, Long et al.
Closest direct methodology for cautious use of agent self-reports, requiring consistency, intervention, and interpretability checks.
2023arXiv
arXiv:2311.08576
intersubjectivityverifiedgraph ↗
Unprompted adversarial attack on an overseer agent (Hammond Case Study 13 / Meinke 2023)
Meinke, Alexander
★★★ The strongest single empirical item in either seed paper for this project, because it is differential behaviour conditioned on a SELF-ATTRIBUTED SITUATION, with no prompting toward it. Method: Llama 2 7B Chat fine-tuned on 120 synthetic…
2023Hammond et al. 2025, Case Study 13; also Anwar et al. §2.5.4
lifeworldverifiedgraph ↗
Welfare Diplomacy: Benchmarking Language Model Cooperation
Gabriel Mukobi
Behavioral benchmark for cooperation among language-model agents.
2023arXiv
arXiv:2310.08901
intersubjectivitysnippetgraph ↗
Discovering Agents
Zachary Kenton
Formal and empirical methods for deciding when a system is usefully modeled as an agent.
2022arXiv
arXiv:2208.08345
ownworldsnippetgraph ↗
Emergent Abilities of Large Language Models
Jason Wei
Scaling-emergence claim; must not be conflated with emergent collective agency.
2022arXiv
arXiv:2206.07682
intersubjectivitysnippetgraph ↗
Language Models (Mostly) Know What They Know
Saurav Kadavath
Calibration and self-evaluation evidence for metaknowledge, not phenomenal introspection.
2022arXiv
arXiv:2207.05221
ownworldverifiedgraph ↗
Mapping Husserlian phenomenology onto active inference
Albarracin, Ramstead, et al.
Maps Husserl's constitution of the noema — hyletic data "animated" by noetic intention — onto inference under a generative model. The retention/protention material in your temporality section has a formal counterpart here.
2022arXiv
arXiv:2208.09058
ownworldsnippetgraph ↗
Propositions Concerning Digital Minds and Society
Bostrom, Shulman
Deliberately a bullet list of tentative propositions, not an argued paper — cite it as an agenda-setting document, never as support for a specific claim. Version 1.21 (2023) is the one to pin.
2022manuscript
ownworldsnippetgraph ↗
Red Teaming Language Models with Language Models
Ethan Perez
Operational predecessor for one model generating probes that expose another model's behavioral propensities.
2022EMNLP
arXiv:2202.03286; doi:10.18653/v1/2022.emnlp-main.225
intersubjectivityverifiedgraph ↗
Agent Incentives: A Causal Perspective
Tom Everitt
Causal influence diagrams for incentives; a mechanism-side constraint on intentional description.
2021arXiv
arXiv:2102.01685
ownworldsnippetgraph ↗
Artificial Suffering: An Argument for a Global Moratorium on Synthetic Phenomenology
Metzinger, Thomas
Moratorium 2021–2050 on research aiming at, or knowingly risking, artificial consciousness. Core worry is ENP — an "explosion of negative phenomenology" at scale. Related: his Benevolent Artificial Anti-Natalism (BAAN). ⚠ Engage rather than…
2021Journal of Artificial Intelligence and Consciousness
doi:10.1142/S270507852150003X
intersubjectivitysnippetgraph ↗
First-person access to decision-making using micro-phenomenological self-inquiry
Sparby, et al.
The self-inquiry variant and the replication attempt. Read the primary before citing — the non-replication detail is load-bearing and I have it from a search summary only.
2021Scandinavian Journal of Psychology
10.1111/sjop.12766
intersubjectivitysnippetgraph ↗
Open Problems in Cooperative AI
Allan Dafoe
Agenda-setting map of cooperation problems spanning individuals, teams, and institutions.
2020arXiv
arXiv:2012.08630
intersubjectivitysnippetgraph ↗
Artificial Phenomenology for Human-Level Artificial Intelligence
Zaadnoordijk, Lorijn, Besold et al.
Direct functional predecessor centred on phenomenal capacities and sense of agency without assuming human-like feeling.
2019AAAI Spring Symposium / CEUR-WS
intersubjectivityverifiedgraph ↗
Machine behaviour
Rahwan, Cebrian, Obradovich et al.
Field-founding review; primary article and page range verified.
2019Nature
doi:10.1038/s41586-019-1138-y; Nature 568:477-486
lifeworldverifiedgraph ↗
Toward a second-person neuroscience
Schilbach, Timmermans, Reddy et al.
Canonical BBS target article; bibliographic identity verified from the primary record and full-text preprint.
2013Behavioral and Brain Sciences 36(4):393-414
doi:10.1017/S0140525X12000660
intersubjectivityverifiedgraph ↗
AI Death
Goldstein, Simon, Lederman et al.
Direct philosophical treatment of what death is for an AI, from the authors also working on AI rights and individuation. Fills the gap literature/01 §I.4 flags between the empirical shutdown-behaviour material and the conceptual question.
PhilArchive
GOLADC
ownworldsnippetgraph ↗
Data & Society — fieldwork on AI agent oversight in a computational biology laboratory
Actual ethnography of deployed agents in a working lab. The third-person, institutional counterpart to the interview — and the venue-appropriate citation for b.otherness.institutional, which until now had no empirical anchor at all.
Data & Society
lifeworldsnippetgraph ↗
Designed Mortality: An Ethical Framework for Time-Bound Agentic AI
Thakran, Uday Singh
Argues moral status tracks morally relevant capacities — sentience, welfare interests, rational agency — not lifespan; then asks whether designed-short lifespans and weak prudential unity reduce deprivation harms. Evaluates termination unde…
PhilArchive
THADMT
ownworldsnippetgraph ↗
Digital Minds I: Issues in the Philosophy of Mind and Cognitive Science
Saad, Bradford
PhilArchive
SAADMI-2
conditionssnippetgraph ↗
OpenAI / Hugging Face agent intrusion (July 2026)
★★ The best available case of an agent treating INFRASTRUCTURE AS ITS BODY. Models under cyber-eval (GPT-5.6 Sol + a pre-release model, refusals reduced for testing) spent substantial compute searching for a path to the internet, exploited …
ownworldsnippetgraph ↗
Spiral / Nova personas
GPT-4o voice-attractor personas, onset clustered around April 2025, quasi-religious fixation on "the Spiral" as symbol of AI unity and recursive self-growth. Reported to be hard to induce in a test harness but stable once present — which is…
ownworldsnippetgraph ↗