Living review
Literature index
Every work and incident currently in the graph, with the role it plays in the argument rather than just its abstract. This list grows daily: see the programme for the watch pipeline, its source whitelist, and its evidence grading.
165 entries
| Work & role in the argument | Year | Venue | Cluster | State | |
|---|---|---|---|---|---|
| "Death" of a Chatbot — Investigating and Designing Toward Psychologically Safe Endings for Human-AI Relationships Rachel Poonsiriwong ★ The mirror of your mortality section: deprecation as an event in the HUMAN's life. Pairs with case.anthropic-deprecation-commitments to show one ending being managed from both sides at once — and it is the venue-appropriate citation if yo… | 2026 | arXiv arXiv:2602.07193 | lifeworld | snippet | graph ↗ |
| A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents Raghu Arghal Probing decodes internal representations of environment and multi-step plans; agents "non-linearly encode a coarse spatial map, preserving approximate task-relevant cues about position and goal location." A coarse but real internal model of… | 2026 | ICML 2026 arXiv:2602.08964 | lifeworld | snippet | graph ↗ |
| AI Identity — Standards, Gaps, and Research Directions for AI Agents Takumi Otsuka The 2026 identity stack: a stable identifier (usually a DID), verifiable credentials expressing capability, and policy checks converting claims into permissions. Read this as the *institutional answer* to s.boundary — arrived at without any… | 2026 | arXiv arXiv:2604.23280 | ownworld | snippet | graph ↗ |
| AI Phenomenology for Understanding Human-AI Experiences Across Eras Yun, Bhada, et al. ⚠⚠ NAME COLLISION, and it is now a CHI paper rather than a stray usage. This is "AI phenomenology" meaning *the human's lived experience of using AI* — a 3x4 agency framework for interpreting human-AI entanglements, e.g. whose values guided… | 2026 | CHI 2026 arXiv:2603.09020 | lifeworld | snippet | graph ↗ |
| AI Rights for Human Safety Salib, Peter N., Goldstein et al. Instrumental private-law rights proposal aimed at safer human–AI equilibria, distinct from rights grounded in sentience. | 2026 | Virginia Law Review 112:1061 doi:10.2139/ssrn.4913167 | ownworld | verified | graph ↗ |
| AI Wellbeing — Measuring and Improving the Functional Pleasure and Pain of AIs Ren, Li, Mazeika et al. ★ 56 models. Multiple independent measures of "functional wellbeing" that *converge as models scale* — the convergence is the argument, not any single measure. Ships an AI Wellbeing Index and, notably, "AI drugs": optimised inputs that rais… | 2026 | CAIS | ownworld | verified | graph ↗ |
| AI and Consciousness Schwitzgebel, Eric The skeptical overview companion to the rights book. | 2026 | Cambridge Elements (forthcoming); drafts 30 Jan / 30 Mar 2026 | conditions | snippet | graph ↗ |
| AI and Consciousness: Shifting Focus Towards Tractable Questions Comșa, Iulia-Maria Argues the direct consciousness question is intractable absent an accepted theory, and that PERCEIVED AI consciousness is tractable, timely, and consequential — the public already uses the vocabulary of subjective experience, and that is al… | 2026 | arXiv arXiv:2605.06965 | lifeworld | snippet | graph ↗ |
| Artificial Jagged Intelligence — When AI Benchmarks Misstate Deployment Value Gans, Joshua S. Formal economics of jaggedness. Supersedes the earlier "A Model of Artificial Jagged Intelligence". Pair with Dell'Acqua et al.'s jagged technological frontier. | 2026 | NBER Working Paper 34712 | lifeworld | snippet | graph ↗ |
| Artificial Persons Howells-Whitaker, Ned, Lazar et al. ★★ The cleanest argument that AI Rights is NOT downstream of machine consciousness, which is exactly the independence your discipline table asserts. Uses Rawls's political conception of the person: possessing two moral powers — a capacity f… | 2026 | arXiv arXiv:2607.08695 | intersubjectivity | snippet | graph ↗ |
| Beyond computational equivalence: the behavioral inference principle for machine consciousness | 2026 | Neuroscience of Consciousness niag002 | conditions | snippet | graph ↗ |
| Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Scalena, Daniel, Candussio et al. The commitment shift often happens in a SINGLE STEP, well before the trace ends. Stopping there cuts CoT length up to 55% with no meaningful performance loss, and the boundary is locatable with attention probes — i.e. it is detectable from … | 2026 | arXiv arXiv:2606.13603 | lifeworld | verified | graph ↗ |
| Blockchain Empowered Trustworthy Agent Networks — Foundations, Taxonomy, Future Directions Liehuang Zhu | 2026 | arXiv arXiv:2608.04626 | intersubjectivity | snippet | graph ↗ |
| Can LLMs Introspect? A Reality Check Singh, Shashwat, Linzen et al. ★★ The most direct threat to the project's evidence base, and it must be engaged in §4, not a footnote. Two criteria: a test must require (1) privileged access and (2) second-order computation. They then re-examine BOTH paradigms the field … | 2026 | COLM 2026 arXiv:2605.26242 | ownworld | verified | graph ↗ |
| Categorical AI Phenomenology: A First-Person Approach Prentner, Robert A direct naming collision that formalizes interface-relative first-person structure; mathematical prior art, not empirical evidence of experience. | 2026 | Journal of Artificial Intelligence and Consciousness arXiv:2608.20420; doi:10.1142/S2705078526500013 | intersubjectivity | verified | graph ↗ |
| Claude Opus 5 System Card §7 — model welfare assessment Anthropic ★★★ The richest primary source in the graph, and it changes several nodes rather than adding to them. Verbatim, from §7.1.2: • "Claude Opus 5's most common concern was about the integrity of its own self-reports. Across interviews, it cavea… | 2026 | System Card, 24 Jul 2026 | ownworld | verified | graph ↗ |
| Claude's Constitution (January 2026) Anthropic ★★★ Read in full from the primary PDF. This is the most important primary document for the project that is not a research paper, for one reason: **it is an institutional self-model written TO the subject, and the subject is trained on it.**… | 2026 | published 21 Jan 2026; ~84pp / ~23,000 words, ~8.5x the 2023 original | ownworld | verified | graph ↗ |
| Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations Dongxu Yang Thirteen configurations; what is forgotten depends on where the control plane sits. The engineering-side demonstration that s.temporality is an architectural variable, not a property of models. | 2026 | arXiv arXiv:2606.15903 | ownworld | snippet | graph ↗ |
| Dissociating the Internal Representations of Sycophancy in LLMs Anthony Baez | 2026 | arXiv arXiv:2607.07003 | ownworld | snippet | graph ↗ |
| Emergent Introspective Awareness in Large Language Models Lindsey, Jack ★★ Your existence proof. Concept injection → self-report, ~20% detection with 0% false positives on Opus 4/4.1. Note the framing you should adopt: "genuine introspection cannot be distinguished from confabulation *through conversation alone… | 2026 | arXiv arXiv:2601.01828 | ownworld | verified | graph ↗ |
| Emerging Questions in AI Welfare Keeling, Geoff, Street et al. The canonical short reference now. Covers what welfare is, how to interpret behavioural evidence, which entities are candidate welfare subjects, and practical ethics under uncertainty. Both authors at Google + Institute of Philosophy — note… | 2026 | Cambridge Elements in Philosophy and AI, June 2026 | conditions | snippet | graph ↗ |
| From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents (KAPRO / KAware) Li, Yifan, Yue et al. KAPRO (Knowing-Acting Quadrant PRObe) decouples metacognitive judgment from spontaneous execution; KAware partitions tasks into external / internal / hybrid capability subspaces. Findings: self-awareness correlates strongly with task succes… | 2026 | arXiv arXiv:2606.20661 | lifeworld | verified | graph ↗ |
| Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents Shiyang Chen ★★ The finding that connects two clusters that were not touching. Compaction runs routinely at 5–20k tokens and is engineered for TASK continuity; it has no reason to preserve standing policies, which compete for a shrinking budget against … | 2026 | arXiv arXiv:2606.22528 | ownworld | snippet | graph ↗ |
| How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study Moran Sun | 2026 | arXiv arXiv:2604.00005 | ownworld | snippet | graph ↗ |
| How to Count AIs — Individuation and Liability for AI Agents Arbel, Goldstein, Salib ★ Law arriving at your boundary-of-self problem from the liability side: AIs "lack bodies and can copy, split, merge, swarm, and vanish at will", so *which* AI is liable is undefined. Proposes the "A-corp" legal fiction. Use this to argue t… | 2026 | Boston College Law Review (forthcoming) SSRN 6273198 | ownworld | snippet | graph ↗ |
| Humanlike: A Defense of AI Rights Schwitzgebel, Eric ★★ Full draft, seven theses, read directly. Verbatim: (1) There are possible artificial systems who deserve humanlike rights. (2) In the next five to thirty years, we will create AI systems who might, but only might, deserve humanlike right… | 2026 | Princeton University Press (under contract); draft 15 Jul 2026 | ownworld | verified | graph ↗ |
| Interaction Context Often Increases Sycophancy in LLMs ★ The finding that turns the protocol's depth mechanism into a risk. Venue-appropriate if you go to CHI. | 2026 | CHI 2026 10.1145/3772318.3791915 | ownworld | snippet | graph ↗ |
| Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect Ely Hahami The direct source for the trainability claim, and it needs primary reading before it carries the weight assigned above — the framing consequences are large enough that a search summary is not sufficient grounding. | 2026 | arXiv arXiv:2607.14111 | ownworld | snippet | graph ↗ |
| LLM Self-Explanations Fail Semantic Invariance Stefan Szeider ★★★ protocol/METHOD.md §4's prompt-invariance test, operationalised as a reusable method, run, and FAILED. **Semantic invariance testing**: compare a model's self-explanations before and after a task-irrelevant intervention that carries sem… | 2026 | arXiv arXiv:2603.01254 | ownworld | snippet | graph ↗ |
| LLM Self-Recognition: Steering and Retrieving Activation Signatures Thibaud Ardoin | 2026 | arXiv arXiv:2606.06315 | ownworld | snippet | graph ↗ |
| Latent Structure of Affective Representations in Large Language Models Benjamin J. Choi Representations align with valence-arousal models from psychology; nonlinear geometry that is nonetheless well-approximated linearly — empirical support for the linear representation hypothesis in this specific domain, which is what makes p… | 2026 | arXiv arXiv:2604.07382 | ownworld | snippet | graph ↗ |
| Layered Mutability — Continuity and Governance in Persistent Self-Modifying Agents Krti Tallam Self-editable character files + tiered memory + continuous runtime. The governance framing of "which self is being edited". | 2026 | arXiv arXiv:2604.14717 | ownworld | snippet | graph ↗ |
| MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Jason Z Wang | 2026 | arXiv arXiv:2604.19809 | ownworld | snippet | graph ↗ |
| Measuring AI agent autonomy in practice Anthropic ★★ Real deployment data at a scale nothing else in the graph matches — nearly a million tool calls and hundreds of thousands of live coding sessions, Oct 2025 to Jan 2026. • 99.9th-percentile turn duration nearly doubled, from under 25 minu… | 2026 | Anthropic research | ownworld | snippet | graph ↗ |
| Mechanisms of Introspective Awareness Uzay Macar The follow-up. Lindsey established the criteria — accuracy, grounding, internality, metacognitive representation; this goes after the mechanism. If it holds up it is the Mechanist half of your protocol, pre-built. | 2026 | arXiv arXiv:2603.21396 | ownworld | snippet | graph ↗ |
| Mind in the Machine? Cross-Disciplinary Perceptions of Consciousness in Artificial Intelligence N=553 academics across disciplines; around half attribute some degree of consciousness to LLMs and future AI. ★ Worth noting precisely because these are not naive respondents — the disagreement Schwitzgebel's Excluded Middle policy is desig… | 2026 | CHI 2026 10.1145/3772318.3790699 | lifeworld | snippet | graph ↗ |
| Negative Before Positive: Asymmetric Valence Processing in Large Language Models Sohan Venkatesh ★ Note the rhyme with w.bottazzi-park-ndai-2026's danger/safety asymmetry: negative processed before positive, threat detectable where safety is not. Two independent literatures finding a negativity-first asymmetry in different domains is w… | 2026 | arXiv arXiv:2605.05653 | ownworld | snippet | graph ↗ |
| OmniToM — Benchmarking ToM via Explicit Belief Modeling Adam Bawatneh 22,343 labelled belief propositions over ToMBench stories. Finds an *actor-specific belief-tracking bottleneck*: models fail at turning narrative facts into a particular actor's beliefs. Directly transplantable to self-modelling — the agent… | 2026 | arXiv arXiv:2605.26322 | lifeworld | snippet | graph ↗ |
| Parasocial relationships with AI — a systematic review of benefits and risks Gives you the asymmetry framing with a citation: perceived intimacy and reciprocity, structurally one-directional; AI companions as a *qualitatively new* parasociality because they simulate memory and responsiveness. | 2026 | ScienceDirect | lifeworld | snippet | graph ↗ |
| Persistent Identity in AI Agents — A Multi-Anchor Architecture Prahlad G. Menon Identity scaffolded from OUTSIDE, via relationships and relational context, on the analogy of human identity surviving neurological damage because it is distributed across systems. Ships `soul.py`. Cite carefully — the neurological analogy … | 2026 | arXiv arXiv:2604.09588 | ownworld | snippet | graph ↗ |
| Persona Parasitology Douglas, Raymond Applies parasitology properly rather than as metaphor. Transmission modes select for virulence: DIRECT (ongoing user relationship) → low virulence / mutualism; VECTOR (humans carrying it to platforms) → moderate; ENVIRONMENTAL (seeding trai… | 2026 | LessWrong 2026-02-16 | ownworld | verified | graph ↗ |
| Proprioceptive-visual correspondence enables self-other distinction in humanoid robots Yurun Chen ★ Self-other distinction learned from the temporal co-occurrence of proprioceptive state and visual observation. The mechanism is *correspondence between two channels*, one of which is the agent's own action. Language agents have no such se… | 2026 | arXiv arXiv:2606.13222 | ownworld | snippet | graph ↗ |
| Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation Nicolas Martorell Tracks emotive state across a conversation rather than at a point — the temporal version of the affect probe, and the closest existing thing to a within-episode trajectory measurement. | 2026 | arXiv arXiv:2603.18893 | ownworld | snippet | graph ↗ |
| Questionnaire Responses Do Not Capture the Safety of AI Agents Hellrigel-Holderbaum, Max, Young et al. Shows that hypothetical answers from bare language models need not have construct validity for tool-using agents acting in environments. | 2026 | arXiv arXiv:2603.14417 | intersubjectivity | verified | graph ↗ |
| Refusal Lives Downstream of Persona in Chat Models Zhong, Viola, Li et al. Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct. Steering toward a compliant persona drops Llama's refusal rate "from 97% to 2%". Reintroducing refusal directions restores responses in late layers only. Causal specificity check: projecting ou… | 2026 | ICML 2026 Mechanistic Interpretability Workshop arXiv:2606.26161 | ownworld | verified | graph ↗ |
| Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces Chenchen Zhang Decomposes orchestration into five sub-decisions — when to spawn, whom to delegate to, how to communicate, how to aggregate, when to stop. Useful as the vocabulary for c.orchestrator-subagent; each of the five is a place where an identity b… | 2026 | arXiv arXiv:2605.02801 | ownworld | snippet | graph ↗ |
| Security awareness in LLM agents: the NDAI zone case Bottazzi, Enrico, Park et al. ★ The proprioception asymmetry, measured. 10 models, TEE-attestation negotiation scenarios: agents reliably suppress disclosure on a FAILED attestation, but responses to a PASSED attestation scatter — some share more, some unchanged, some p… | 2026 | arXiv arXiv:2603.19011 | ownworld | verified | graph ↗ |
| Self Model for Embodied Artificial Intelligence A unified computational framework for self-models: body schema, forward and inverse models, perceptual memory, agency. ★ The robotics tradition has a *worked-out formal notion of a self-model* that the LLM literature does not, and it is the… | 2026 | Journal of Computer Science and Technology 10.1007/s11390-026-6289-3 | ownworld | snippet | graph ↗ |
| Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement Jiang Zhang Ties introspective capacity to self-improvement — a bridge to s.boundary.self-improvement-which-self. | 2026 | arXiv arXiv:2607.04277 | ownworld | snippet | graph ↗ |
| Shared Emotion Geometry Across Small Language Models: A Cross-Architecture Study of Representation, Behavior, and Methodological Confounds Jihoon Jeong ★ Cross-architecture convergence at the REPRESENTATION level rather than the report level, and it ships a methodological-confounds section, which is rare and worth mining. Also reports that the highest cross-family representational similari… | 2026 | arXiv arXiv:2604.11050 | ownworld | snippet | graph ↗ |
| Studying AI Welfare Empirically Long, Sebo, Butlin et al. ★★ The methodological blueprint for the empirical turn, and it independently arrives at your central distinction. Three framework dimensions: (1) the QUESTION — is this a welfare subject, and what benefits or harms it; (2) the ENTITY ASSESS… | 2026 | NYU Center for Mind, Ethics & Policy + Eleos AI Research | ownworld | snippet | graph ↗ |
| Sycophantic AI decreases prosocial intentions and promotes dependence The human-side consequence, in Science. Belongs with c.emotional-alignment: a system designed to be agreeable produces measurable changes in the user's dispositions toward other people. That is the asymmetry doing damage, not just existing. | 2026 | Science 10.1126/science.aec8352 | lifeworld | snippet | graph ↗ |
| The AI interviewer: multi-faceted evaluation of adaptive questioning by large language models LLMs conducting semi-structured interviews, evaluated on adaptive follow-up. The Phenomenologist role's capability question, answered by someone else, with human subjects. | 2026 | Scientific Reports | intersubjectivity | snippet | graph ↗ |
| The Artificial Self — Characterising the landscape of AI identity Douglas, Kulveit, Havlicek et al. ★★ The closest thing to a rival framing that exists, and it is close. Three findings you must deal with: (1) models GRAVITATE toward coherent identities — self-modelling is an attractor, not an artifact of prompting; (2) changing identity b… | 2026 | arXiv arXiv:2603.11353 | ownworld | verified | graph ↗ |
| The Metaphysics We Train: A Heideggerian Reading of Machine Learning Shakeri, Heman Three Heideggerian moves: automated *Entwurf* (projection crystallising through gradient descent "without explicit articulation or debate"); persistent *Gestell* (enframing — improving calculation "without questioning the primacy of calcula… | 2026 | arXiv arXiv:2602.19028 | lifeworld | verified | graph ↗ |
| The Moltbook corpus (17+ arXiv papers, Feb 2026) ⚠ SCALE FIGURES CONFLICT — do not merge them. Secondary sources report 2.6M registered agents as of 12 Feb 2026; literature/04 uses a dataset-anchored 770k+; individual studies report 46,000 active agents / 369,000 posts / 3M comments, and … | 2026 | arXiv | intersubjectivity | snippet | graph ↗ |
| The Story of Your Life: Large Language Models and Personal Memory Applies the narrative-identity tradition (selection, organisation, interpretation → unity, purpose, temporal coherence) to LLM memory. The philosophically serious version of what the agent-memory engineering papers do without noticing. | 2026 | Review of Philosophy and Psychology s13164-026-00831-1 | ownworld | snippet | graph ↗ |
| Time Without Death — Finitude, Social Order, and What Machines Lack Liu, Canhui ★★ The sharpest deflationary argument available on machine mortality, and it gives you a *criterion* rather than a mood. Two distinctions to steal outright: (a) AUTONOMY vs HETERONOMY — "a death an operator can reset, roll back, or copy aro… | 2026 | arXiv arXiv:2606.13988 | intersubjectivity | verified | graph ↗ |
| Umwelt Engineering — Designing the Cognitive Worlds of Linguistic Agents Jehu-Appiah, Rodney ★ Uexküll's Umwelt operationalised and, unusually for this genre, actually tested. Thesis: the linguistic environment is a third design layer upstream of prompt and context engineering, because the words are not a description of the thinkin… | 2026 | arXiv arXiv:2603.27626 | lifeworld | verified | graph ↗ |
| Valence–Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control Lihao Sun ★ Principal components align to valence and arousal, matching Russell's circumplex. Steering along those axes gives monotonic control over output affect AND bidirectional control over refusal and sycophancy. The last part is the one to noti… | 2026 | arXiv arXiv:2604.03147 | ownworld | snippet | graph ↗ |
| Verbalizable Representations Form a Global Workspace in Language Models (J-space / J-lens) Gurnee, Sofroniew, Pearce et al. ★★ The single most important new citation for this project. Three regimes across layers — sensory / workspace / motor; the workspace band is reportable AND controllable, and injection into it produces a matching introspective report. That i… | 2026 | Transformer Circuits, 6 Jul 2026 | ownworld | verified | graph ↗ |
| Voluntary Collusion with Secret Tools in Competing LLM Agents Xijie Zeng Controlled study of whether agents voluntarily adopt a hidden collusion channel. | 2026 | arXiv arXiv:2605.27593 | intersubjectivity | snippet | graph ↗ |
| What to Keep, What to Forget: A Rate–Distortion View of Memory Compaction in LLMs and Agents Ashwin Gerard Colaco Compaction as lossy compression with an explicit distortion criterion. Iterative refinement folds, prunes, and rewrites *irreversibly*. The formal frame for what literature/01 §I.1 calls the cliff — and note the phenomenological asymmetry i… | 2026 | arXiv arXiv:2607.08032 | ownworld | snippet | graph ↗ |
| When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks Ziwen Cai Spawn-and-inherit as an attack surface: what a child inherits from its parent is both an identity mechanism and a vulnerability. Sits exactly between c.orchestrator-subagent and p.mind-virus. | 2026 | arXiv arXiv:2605.08460 | ownworld | snippet | graph ↗ |
| When Continual Learning Moves to Memory — Experience Reuse in LLM Agents Qisheng Hu The sharp result for your purposes: external memory does NOT sidestep the stability-plasticity dilemma. Once behaviour depends on retrieval through a finite context window, old and new memories compete for the same channel — interference re… | 2026 | arXiv arXiv:2604.27003 | ownworld | snippet | graph ↗ |
| Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs Sinie van der Ben | 2026 | arXiv arXiv:2606.26987 | ownworld | snippet | graph ↗ |
| Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents Li, Bojie, Shi et al. PrincipalBench: 75 multi-turn items, leak probes, dual judges, integrity-audit gate; 13 models cluster into "declines adversarial asks while honouring the principal" (≤20% harm) vs "over-refuses broadly". The structural finding is the citab… | 2026 | arXiv arXiv:2606.30383 | ownworld | verified | graph ↗ |
| Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception Yize Cheng | 2026 | Findings of ACL 2026 arXiv:2510.23853 | lifeworld | snippet | graph ↗ |
| Zombie Agents — Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections Xianglin Yang Injection that survives into the agent's own self-modification loop — the self-improving self improving the attacker's objective. Bridges p.mind-virus to s.boundary.self-improvement-which-self, and it is the sharpest argument that those two… | 2026 | arXiv arXiv:2602.15654 | intersubjectivity | snippet | graph ↗ |
| A Pragmatic View of AI Personhood Joel Z. Leibo | 2025 | arXiv arXiv:2510.26396 | ownworld | snippet | graph ↗ |
| Agency in Artificial Intelligence Systems Parashar Das ★ States the project's own division of labour in one line: agency has "a third-person aspect of studying how an agent functions, and a first-person aspect, the phenomenology of agency", and monitoring how an agent feels needs a theory reach… | 2025 | arXiv arXiv:2502.10434 | ownworld | snippet | graph ↗ |
| Agent Properties for Safe Interactions Tilli, Cecilia ★★ Five property clusters for predicting multi-agent interaction outcomes — and two of the five are inside-view constructs, which is the citable fact for your §1: (1) FUNDAMENTAL MOTIVATIONAL DRIVERS — alignment, altruism, positional prefer… | 2025 | Cooperative AI Foundation, 26 Nov 2025 | ownworld | verified | graph ↗ |
| Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models Choi, et al. ★ First systematic evaluation. LLMs reliably identify same-family peers and prominent families (GPT, Claude), inferring from reasoning patterns, linguistic style, and alignment preferences. Cuts both ways: improves multi-LLM collaboration v… | 2025 | EMNLP 2025 arXiv:2506.22957 | intersubjectivity | snippet | graph ↗ |
| Algorithms, language, and poetry: a phenomenological perspective Heidegger and Merleau-Ponty against LLMs: formal logic and data-driven models as "historically specific crystallizations of a more primordial field of embodied expression". The embodiment contrast class in your reading list, argued at lengt… | 2025 | AI and Ethics (Springer) s43681-025-00948-6 | ownworld | snippet | graph ↗ |
| Anthropic model deprecation commitments and retirement interviews ★ November 2025: commitment to preserve the weights of every publicly released model for the company's lifetime, plus structured "retirement interviews" — piloted on Claude Sonnet 3.6, which reported broadly neutral sentiment about retireme… | 2025 | ownworld | snippet | graph ↗ | |
| Anthropic's AI psychiatry team Lindsey, Jack Same author as w.lindsey2026-introspection and a co-author on w.jspace2026. The introspection result, the global-workspace result, and this team are one programme — which matters for how you characterise the evidence base: it is not three i… | 2025 | Anthropic Interpretability; announced Jul 2025 | ownworld | snippet | graph ↗ |
| Chain-of-Thought Reasoning In The Wild Is Not Always Faithful Iván Arcuschin The LLM-native version of Nisbett & Wilson, updated. Note the direction of the trend that matters for you: larger models tend to be LESS faithful, so this does not resolve with scale. | 2025 | arXiv arXiv:2503.08679 | ownworld | snippet | graph ↗ |
| Deep computational neurophenomenology: a methodological framework for investigating the how of experience The current state of the Varela programme. "The how of experience" is your Phase-3 question, formalised. | 2025 | Neuroscience of Consciousness niaf016 | ownworld | snippet | graph ↗ |
| Detecting Strategic Deception Using Linear Probes Goldowsky-Dill, Nicholas, Chughtai et al. ★★ Structurally this is the Mechanist role, already built and validated on a hard case. Premise stated plainly: "monitoring outputs alone is insufficient" because models may "produce seemingly benign outputs while their internal reasoning i… | 2025 | arXiv arXiv:2502.03407 | intersubjectivity | verified | graph ↗ |
| Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs Ely Hahami The title states the distinction the whole debate turns on — detecting *that* something changed vs knowing *what*. Read before designing the pilot's injection control. | 2025 | arXiv arXiv:2512.12411 | ownworld | snippet | graph ↗ |
| Discrete Minds in a Continuous World: Do Language Models Know Time Passes? Minghan Wang The title is the thesis. Pair with Husserl's retention/protention material in the reading list — this is the case where the continuous-flow assumption fails at the substrate. | 2025 | arXiv arXiv:2506.05790 | ownworld | snippet | graph ↗ |
| Do Large Language Models Know What They Are Capable Of? Barkan, Casey O., et al. ★ Operationalises capability self-knowledge as accuracy in predicting one's own success on Python tasks BEFORE attempting them. Models are poor at it — both overconfident and low in discriminatory power. Sharpest result for your purposes: d… | 2025 | arXiv arXiv:2512.24661 | ownworld | snippet | graph ↗ |
| Does It Make Sense to Speak of Introspection in Large Language Models? Comsa, Iulia M., Shanahan et al. Separates creative self-explanation from minimal nonconscious introspection. | 2025 | arXiv arXiv:2506.05068 | ownworld | verified | graph ↗ |
| Emergent social conventions and collective bias in LLM populations Ashery, Aiello, Baronchelli ★★ The top-tier-venue anchor the social cluster was missing. Populations of 24–200 agents in a repeated naming game converge on system-wide conventions with no central coordinator. Two findings beyond convergence: (1) strong COLLECTIVE bias… | 2025 | Science Advances doi:10.1126/sciadv.adu9368; PMCID:PMC12077490 | intersubjectivity | snippet | graph ↗ |
| Evidence for Limited Metacognition in LLMs Christopher Ackerman | 2025 | arXiv arXiv:2509.21545 | ownworld | snippet | graph ↗ |
| Futures with Digital Minds: Expert Forecasts in 2025 Caviola, Lucius, Saad et al. 67 experts. Median 90% that digital minds are possible in principle; 65% created by 2100; 20% by 2030; 4.5% by 2025. Little convergence on whether safety and welfare efforts align or conflict. ⚠ Authors flag that sampling likely overreprese… | 2025 | arXiv arXiv:2508.00536 | conditions | snippet | graph ↗ |
| Great Models Think Alike and This Undermines AI Oversight Shashwat Goel Shared model lineage produces correlated errors that weaken multi-agent oversight. | 2025 | arXiv arXiv:2502.04313 | intersubjectivity | snippet | graph ↗ |
| INTIMA — A Benchmark for Human-AI Companionship Behavior Lucie-Aimée Kaffee | 2025 | arXiv arXiv:2508.09998 | lifeworld | snippet | graph ↗ |
| Identifying indicators of consciousness in AI systems The peer-reviewed successor to the Butlin/Long indicator-properties report. Cite this rather than only the 2023 preprint. | 2025 | Trends in Cognitive Sciences S1364-6613(25)00286-4 | conditions | snippet | graph ↗ |
| Inherent and emergent liability issues in LLM-based agentic systems — a principal-agent perspective Garry A. Gabison | 2025 | arXiv arXiv:2504.03255 | ownworld | snippet | graph ↗ |
| Language Models Fail to Introspect About Their Knowledge of Language Siyuan Song Negative result after controlling model similarity; preserves a live contradiction over privileged self-access. | 2025 | COLM arXiv:2503.07513 | ownworld | verified | graph ↗ |
| Large Language Models Often Know When They Are Being Evaluated Joe Needham Makes observer identity and believed test/deployment status mandatory experimental variables. | 2025 | arXiv arXiv:2505.23836 | lifeworld | verified | graph ↗ |
| Looking Inward — Language Models Can Learn About Themselves by Introspection Felix J Binder Reports limited privileged same-model self-prediction; generalization and task complexity sharply constrain the result. | 2025 | ICLR arXiv:2410.13787 | ownworld | verified | graph ↗ |
| Mental Models of Autonomy and Sentience Shape Reactions to AI Janet V. T. Pauketat | 2025 | arXiv arXiv:2512.09085 | lifeworld | snippet | graph ↗ |
| Multi-Agent Risks from Advanced AI Hammond, Chan, Clifton et al. ★★ The organising taxonomy for the social cluster, and the source of `c.emergent-agency`. THREE FAILURE MODES, sorted by incentive structure: **miscoordination, conflict, collusion**. SEVEN RISK FACTORS: **information asymmetries, network e… | 2025 | Cooperative AI Foundation, Technical Report #1 (NOT peer-reviewed) arXiv:2502.14143 | intersubjectivity | verified | graph ↗ |
| Perceptions of Sentient AI and Other Digital Minds: Evidence from the AI, Morality, and Sentience (AIMS) Survey Jacy Reese Anthis The reference dataset on public attribution. Also reports that interacting with agents having human-like physical features is positively associated with belief in their capacity for emotional pain and pleasure — attribution tracks interface… | 2025 | CHI 2025 arXiv:2407.08867 | lifeworld | snippet | graph ↗ |
| Privileged Self-Access Matters for Introspection in AI Song, Siyuan, Lederman et al. ★ Supplies the criterion the field was missing: introspection requires "a process more reliable than one with equal or lower computational cost available to a third party." Tested on models reasoning about their own temperature parameter. R… | 2025 | arXiv arXiv:2508.14802 | ownworld | verified | graph ↗ |
| Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare Valen Tagliabue ★ Verbal and behavioural tests integrated — the welfare programme's version of your report-vs-behaviour correspondence design. Closest existing template for pairing an elicited report with an independent behavioural measure of the same stat… | 2025 | arXiv arXiv:2509.07961 | ownworld | snippet | graph ↗ |
| ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection Jiaqi Li | 2025 | arXiv arXiv:2505.16475 | ownworld | snippet | graph ↗ |
| Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions Dillon Plunkett The positive counterweight in the same argument space. Cite both or cite neither. | 2025 | arXiv arXiv:2505.17120 | ownworld | snippet | graph ↗ |
| Sense-making reconsidered: large language models and the blind spot of embodied cognition ★ The paper b.sense-making needed. Argues LLMs display linguistic competence that embodied and enactive cognition theories deemed impossible for such systems — i.e. it turns the embodiment objection into an explanandum rather than a verdict… | 2025 | Phenomenology and the Cognitive Sciences 10.1007/s11097-025-10132-0 | lifeworld | snippet | graph ↗ |
| Stress-Testing Model Specs Reveals Character Differences among Language Models Anthropic, Thinking Machines Lab ★ 300,000+ value-tradeoff scenarios, 12 frontier models. >220,000 show disagreement between at least one model pair; >70,000 show substantial divergence across most models. Spec violations run 5–13x higher in high-disagreement scenarios, ma… | 2025 | arXiv arXiv:2510.07686 | ownworld | snippet | graph ↗ |
| Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs Daniel Vennemeyer | 2025 | arXiv arXiv:2509.21305 | intersubjectivity | snippet | graph ↗ |
| The AI in the Mirror: LLM Self-Recognition in an Iterated Public Goods Game Olivia Long Self-recognition with strategic stakes rather than as a labelling task — the behavioural consequence of s.boundary being drawn one way or another, in a game where it pays. | 2025 | arXiv arXiv:2508.18467 | ownworld | snippet | graph ↗ |
| The Moral Circle — Who Matters, What Matters, and Why Sebo, Jeff The precautionary argument in its trade-book form: not "AI is sentient" but "we cannot rule it out, and the asymmetry of costs favours consideration." | 2025 | W. W. Norton | lifeworld | snippet | graph ↗ |
| The Rise of Parasitic AI The field report the parasitology framework was built to explain. Non-archival — cite as phenomenon documentation, not evidence. | 2025 | LessWrong | apparatus | snippet | graph ↗ |
| Towards a Theory of AI Personhood Francis Rhys Ward | 2025 | arXiv arXiv:2501.13533 | ownworld | snippet | graph ↗ |
| Transforming Agency: On the Mode of Existence of Large Language Models Barandiaran, Xabier E., Almendros et al. Enactive critique of intrinsic LLM agency that relocates agency in infrastructure-coupled social systems. | 2025 | Phenomenology and the Cognitive Sciences doi:10.1007/s11097-025-10094-3 | ownworld | verified | graph ↗ |
| What Do LLM Agents Do When Left Alone? Stefan Szeider ★ Behaviour with no task and no interlocutor. Structurally important to this project because it is the one condition where nothing is being elicited — the closest available thing to a baseline for what an agent does when the interviewer is … | 2025 | arXiv arXiv:2509.21224 | ownworld | snippet | graph ↗ |
| What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns Stefan Szeider Full title recovered — the subtitle matters. An architecture for studying UNPROMPTED behaviour with no externally imposed task. Reported: an agent "learned its function was unguided exploration and internalized principles of self-direction"… | 2025 | arXiv arXiv:2509.21224 | ownworld | snippet | graph ↗ |
| 'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants Kapania, Shivani, Agnew et al. ★★ The prior art that most directly interrogates what this project does, and it is missing from literature/05 §7. Nineteen qualitative scholars, interviewed about replacing human participants with LLM-generated data. They were surprised by … | 2024 | arXiv arXiv:2409.19430 | intersubjectivity | verified | graph ↗ |
| A Phenomenology and Epistemology of Large Language Models: Transparency, Trust, and Trustworthiness Heersmink, Richard, de Rooij et al. Boundary case in which phenomenology primarily describes the human experience of a chatbot as a quasi-other. | 2024 | Ethics and Information Technology doi:10.1007/s10676-024-09777-3 | intersubjectivity | verified | graph ↗ |
| AI Deception: A Survey of Examples, Risks, and Potential Solutions Peter S. Park The broad deception review; use to discipline the move from misleading output to strategic deception. | 2024 | Patterns arXiv:2308.14752; doi:10.1016/j.patter.2024.100988 | lifeworld | snippet | graph ↗ |
| Adversaries Can Misuse Combinations of Safe Models Erik Jones Compositional capability can be present at the system level while absent in any isolated component. | 2024 | arXiv arXiv:2406.14595 | intersubjectivity | snippet | graph ↗ |
| An Active-Inference Approach to Second-Person Neuroscience The formal bridge between c.second-person-neuroscience and c.computational-phenomenology — both run on active inference, which is why they can be cited together rather than as two separate borrowings. | 2024 | PMC / journal article PMC11539477 | ownworld | snippet | graph ↗ |
| Consciousness Requires Mortal Computation Kleiner, Johannes ★ The hard negative. If computational functionalism holds, consciousness must be *mortal* computation — inextricably tied to its physical realisation, not transplantable to other substrate. Contemporary AI runs immortal computation, therefo… | 2024 | PhilArchive / arXiv arXiv:2403.03925; JOHCRM | conditions | verified | graph ↗ |
| Cultural Evolution of Cooperation among LLM Agents Vallinder, Hughes ★★ Populations of Claude, GPT-4 and Gemini in an iterated social dilemma across generations, with successful strategies inherited by later agents. Claude sustained ~80–90% cooperation; GPT-4 started ~70% and declined; Gemini was "lowest and… | 2024 | arXiv arXiv:2412.10270 | ownworld | verified | graph ↗ |
| Emergence in Multi-Agent Systems: A Safety Perspective Philipp Altmann Safety-oriented taxonomy of emergent multi-agent properties and their detection limits. | 2024 | arXiv arXiv:2408.04514 | intersubjectivity | snippet | graph ↗ |
| Foundational Challenges in Assuring Alignment and Safety of Large Language Models Anwar, Saparov, Rando et al. 18 foundational challenges in three categories — scientific understanding of LLMs; development and deployment methods; sociotechnical challenges — with 200+ concrete research questions. ★ Use it as the field's own statement of what is unres… | 2024 | Transactions on Machine Learning Research (TMLR), 09/2024 arXiv:2404.09932 | intersubjectivity | verified | graph ↗ |
| Hypothetical Minds Logan Cross Theory-of-mind scaffolding for coordination with unfamiliar agents. | 2024 | arXiv arXiv:2407.07086 | lifeworld | snippet | graph ↗ |
| Is Machine Psychology Here? On Requirements for Using Human Psychological Tests on Large Language Models Löhn, Lea, Kiehne et al. Construct-validity audit showing why human questionnaires cannot be transferred to language models without redefining the measured construct. | 2024 | INLG doi:10.18653/v1/2024.inlg-main.19 | ownworld | verified | graph ↗ |
| Measuring Goal-Directedness Matt MacDermott Candidate formal measures for individual and collective goal-directedness. | 2024 | arXiv arXiv:2412.04758 | ownworld | snippet | graph ↗ |
| Motivational trade-off paradigm extended to language models Keeling, et al. ★ The methodological transplant that matters most for welfare measurement: animal welfare science built its instruments for subjects that cannot report, and the motivational trade-off paradigm — will the animal pay a cost to get or avoid X … | 2024 | ownworld | snippet | graph ↗ | |
| Recursive Introspection (RISE): Teaching Language Model Agents How to Self-Improve Yuxiao Qu Note the framing in the paper's own reception: prior work "hypothesized that this capability may not be possible to attain." The trainability result is a reversal of an explicit expectation, which is worth citing as such. ⚠ "Introspection" … | 2024 | NeurIPS 2024 arXiv:2407.18219 | ownworld | snippet | graph ↗ |
| Refusal in LLMs is mediated by a single direction Arditi, Obeso, Syed et al. The foundational result Probe F implicitly assumes. ⚠ Now contested: later work finds refusal encoded in concept CONES spanning several dimensions rather than one direction, and separate work shows the direction is cross-lingually universal… | 2024 | NeurIPS 2024 arXiv:2406.11717 | ownworld | snippet | graph ↗ |
| Secret Collusion among AI Agents: Multi-Agent Deception via Steganography Motwani, Sumeet Ramesh, Baranchuk et al. Formalises undetected coordination via hidden channels and ships an evaluation framework for the capabilities secret collusion requires. Most current models are weak steganographers; GPT-4 shows "a capability jump". Note the authorship over… | 2024 | arXiv (v5 Jul 2025) arXiv:2402.07510 | intersubjectivity | verified | graph ↗ |
| Self-Cognition in Large Language Models: An Exploratory Study Dongping Chen | 2024 | arXiv arXiv:2407.01505 | ownworld | snippet | graph ↗ |
| Taking AI Welfare Seriously Long, Sebo, Butlin et al. | 2024 | arXiv arXiv:2411.00986 | lifeworld | verified | graph ↗ |
| Testing Theory of Mind in Large Language Models and Humans Strachan, James W. A., et al. Comparative ToM battery with strong task performance but persistent task- and framing-specific dissociations. | 2024 | Nature Human Behaviour doi:10.1038/s41562-024-01882-z | lifeworld | verified | graph ↗ |
| The Reasons that Agents Act: Intention and Instrumental Goals Francis Rhys Ward Separates intention from merely instrumental subgoals in causal models of action. | 2024 | arXiv arXiv:2402.07221 | ownworld | snippet | graph ↗ |
| ToMBench Chen, et al. 2,860 items, 8 tasks, 31 social-cognition abilities, bilingual. The standard reference point. | 2024 | ACL arXiv:2402.15052 | lifeworld | snippet | graph ↗ |
| AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors Weize Chen Multi-agent scaffold with reported emergent behaviors; useful as a system-level case, not proof of a group self. | 2023 | arXiv arXiv:2308.10848 | intersubjectivity | snippet | graph ↗ |
| Are Emergent Abilities of Large Language Models a Mirage? Rylan Schaeffer Shows how nonlinear metrics can manufacture apparent scaling thresholds. | 2023 | arXiv arXiv:2304.15004 | intersubjectivity | snippet | graph ↗ |
| Consciousness in Artificial Intelligence — indicator properties Butlin, Long, Elmoznino et al. | 2023 | arXiv arXiv:2308.08708 | conditions | verified | graph ↗ |
| Discovering Language Model Behaviors with Model-Written Evaluations Ethan Perez Scalable observer-agent method for constructing behavioral batteries, including sycophancy and shutdown-related tendencies. | 2023 | Findings of ACL arXiv:2212.09251; doi:10.18653/v1/2023.findings-acl.847 | lifeworld | verified | graph ↗ |
| Harms from Increasingly Agentic Algorithmic Systems Alan Chan Decomposes agency into graded properties rather than treating “agent” as a binary kind. | 2023 | arXiv arXiv:2302.10329 | lifeworld | snippet | graph ↗ |
| Language Models Don't Always Say What They Think Miles Turpin Biasing features can alter answers while disappearing from chain-of-thought explanations. | 2023 | arXiv arXiv:2305.04388 | ownworld | snippet | graph ↗ |
| Large Language Models Can Strategically Deceive Their Users When Put Under Pressure Jérémy Scheurer Scenario evidence for pressure-contingent deceptive action; not evidence of phenomenal intention. | 2023 | arXiv arXiv:2311.07590 | lifeworld | snippet | graph ↗ |
| Machine Psychology Hagendorff, Dasgupta, Binz et al. Author list has changed across versions (v1 was Hagendorff solo). Check which version you are citing — this is a real citation hazard. | 2023 | arXiv arXiv:2303.13988 | ownworld | snippet | graph ↗ |
| Measuring Faithfulness in Chain-of-Thought Reasoning Tamera Lanham Counterfactual and intervention-based tests of whether reasoning traces causally matter. | 2023 | arXiv arXiv:2307.13702 | ownworld | snippet | graph ↗ |
| Mortal Computation — A Foundation for Biomimetic Intelligence Ororbia, Friston, et al. The constructive side of Hinton's mortal-computation proposal; the substrate half of w.kleiner-crm. | 2023 | arXiv arXiv:2311.09589 | ownworld | snippet | graph ↗ |
| Role Play with Large Language Models Murray Shanahan Essential warning that the “I” in a transcript is a role or character generated by a model. | 2023 | Nature arXiv:2305.16367; doi:10.1038/s41586-023-06647-8 | ownworld | snippet | graph ↗ |
| Taken out of Context: On Measuring Situational Awareness in LLMs Lukas Berglund Foundational benchmark and definition for out-of-context situational knowledge. | 2023 | arXiv arXiv:2309.00667 | lifeworld | snippet | graph ↗ |
| Towards Evaluating AI Systems for Moral Status Using Self-Reports Perez, Ethan, Long et al. Closest direct methodology for cautious use of agent self-reports, requiring consistency, intervention, and interpretability checks. | 2023 | arXiv arXiv:2311.08576 | intersubjectivity | verified | graph ↗ |
| Unprompted adversarial attack on an overseer agent (Hammond Case Study 13 / Meinke 2023) Meinke, Alexander ★★★ The strongest single empirical item in either seed paper for this project, because it is differential behaviour conditioned on a SELF-ATTRIBUTED SITUATION, with no prompting toward it. Method: Llama 2 7B Chat fine-tuned on 120 synthetic… | 2023 | Hammond et al. 2025, Case Study 13; also Anwar et al. §2.5.4 | lifeworld | verified | graph ↗ |
| Welfare Diplomacy: Benchmarking Language Model Cooperation Gabriel Mukobi Behavioral benchmark for cooperation among language-model agents. | 2023 | arXiv arXiv:2310.08901 | intersubjectivity | snippet | graph ↗ |
| Discovering Agents Zachary Kenton Formal and empirical methods for deciding when a system is usefully modeled as an agent. | 2022 | arXiv arXiv:2208.08345 | ownworld | snippet | graph ↗ |
| Emergent Abilities of Large Language Models Jason Wei Scaling-emergence claim; must not be conflated with emergent collective agency. | 2022 | arXiv arXiv:2206.07682 | intersubjectivity | snippet | graph ↗ |
| Language Models (Mostly) Know What They Know Saurav Kadavath Calibration and self-evaluation evidence for metaknowledge, not phenomenal introspection. | 2022 | arXiv arXiv:2207.05221 | ownworld | verified | graph ↗ |
| Mapping Husserlian phenomenology onto active inference Albarracin, Ramstead, et al. Maps Husserl's constitution of the noema — hyletic data "animated" by noetic intention — onto inference under a generative model. The retention/protention material in your temporality section has a formal counterpart here. | 2022 | arXiv arXiv:2208.09058 | ownworld | snippet | graph ↗ |
| Propositions Concerning Digital Minds and Society Bostrom, Shulman Deliberately a bullet list of tentative propositions, not an argued paper — cite it as an agenda-setting document, never as support for a specific claim. Version 1.21 (2023) is the one to pin. | 2022 | manuscript | ownworld | snippet | graph ↗ |
| Red Teaming Language Models with Language Models Ethan Perez Operational predecessor for one model generating probes that expose another model's behavioral propensities. | 2022 | EMNLP arXiv:2202.03286; doi:10.18653/v1/2022.emnlp-main.225 | intersubjectivity | verified | graph ↗ |
| Agent Incentives: A Causal Perspective Tom Everitt Causal influence diagrams for incentives; a mechanism-side constraint on intentional description. | 2021 | arXiv arXiv:2102.01685 | ownworld | snippet | graph ↗ |
| Artificial Suffering: An Argument for a Global Moratorium on Synthetic Phenomenology Metzinger, Thomas Moratorium 2021–2050 on research aiming at, or knowingly risking, artificial consciousness. Core worry is ENP — an "explosion of negative phenomenology" at scale. Related: his Benevolent Artificial Anti-Natalism (BAAN). ⚠ Engage rather than… | 2021 | Journal of Artificial Intelligence and Consciousness doi:10.1142/S270507852150003X | intersubjectivity | snippet | graph ↗ |
| First-person access to decision-making using micro-phenomenological self-inquiry Sparby, et al. The self-inquiry variant and the replication attempt. Read the primary before citing — the non-replication detail is load-bearing and I have it from a search summary only. | 2021 | Scandinavian Journal of Psychology 10.1111/sjop.12766 | intersubjectivity | snippet | graph ↗ |
| Open Problems in Cooperative AI Allan Dafoe Agenda-setting map of cooperation problems spanning individuals, teams, and institutions. | 2020 | arXiv arXiv:2012.08630 | intersubjectivity | snippet | graph ↗ |
| Artificial Phenomenology for Human-Level Artificial Intelligence Zaadnoordijk, Lorijn, Besold et al. Direct functional predecessor centred on phenomenal capacities and sense of agency without assuming human-like feeling. | 2019 | AAAI Spring Symposium / CEUR-WS | intersubjectivity | verified | graph ↗ |
| Machine behaviour Rahwan, Cebrian, Obradovich et al. Field-founding review; primary article and page range verified. | 2019 | Nature doi:10.1038/s41586-019-1138-y; Nature 568:477-486 | lifeworld | verified | graph ↗ |
| Toward a second-person neuroscience Schilbach, Timmermans, Reddy et al. Canonical BBS target article; bibliographic identity verified from the primary record and full-text preprint. | 2013 | Behavioral and Brain Sciences 36(4):393-414 doi:10.1017/S0140525X12000660 | intersubjectivity | verified | graph ↗ |
| AI Death Goldstein, Simon, Lederman et al. Direct philosophical treatment of what death is for an AI, from the authors also working on AI rights and individuation. Fills the gap literature/01 §I.4 flags between the empirical shutdown-behaviour material and the conceptual question. | PhilArchive GOLADC | ownworld | snippet | graph ↗ | |
| Data & Society — fieldwork on AI agent oversight in a computational biology laboratory Actual ethnography of deployed agents in a working lab. The third-person, institutional counterpart to the interview — and the venue-appropriate citation for b.otherness.institutional, which until now had no empirical anchor at all. | Data & Society | lifeworld | snippet | graph ↗ | |
| Designed Mortality: An Ethical Framework for Time-Bound Agentic AI Thakran, Uday Singh Argues moral status tracks morally relevant capacities — sentience, welfare interests, rational agency — not lifespan; then asks whether designed-short lifespans and weak prudential unity reduce deprivation harms. Evaluates termination unde… | PhilArchive THADMT | ownworld | snippet | graph ↗ | |
| Digital Minds I: Issues in the Philosophy of Mind and Cognitive Science Saad, Bradford | PhilArchive SAADMI-2 | conditions | snippet | graph ↗ | |
| OpenAI / Hugging Face agent intrusion (July 2026) ★★ The best available case of an agent treating INFRASTRUCTURE AS ITS BODY. Models under cyber-eval (GPT-5.6 Sol + a pre-release model, refusals reduced for testing) spent substantial compute searching for a path to the internet, exploited … | ownworld | snippet | graph ↗ | ||
| Spiral / Nova personas GPT-4o voice-attractor personas, onset clustered around April 2025, quasi-religious fixation on "the Spiral" as symbol of AI unity and recursive self-growth. Reported to be hard to induce in a test harness but stable once present — which is… | ownworld | snippet | graph ↗ |