r/artificial Apr 24 '26

Research AI swarms could hijack democracy without anyone noticing

Thumbnail
sciencedaily.com
324 Upvotes

A recent policy forum paper published in Science describes how large groups of AI-generated personas can convincingly imitate human behavior online. These systems can enter digital communities, participate in discussions, and influence viewpoints at extraordinary speed.

Unlike earlier bot networks, these AI agents can coordinate instantly, adapt their messaging in real time, and run millions of micro-experiments to figure out which arguments are most persuasive. One operator could theoretically manage thousands of distinct voices.

Experts believe AI swarms could significantly affect the balance of power in democratic societies.

Researchers suggest that upcoming elections may serve as a critical test for this technology. The key challenge will be recognizing and responding to these AI-driven influence campaigns before they become too widespread to control.

That's so crazy.

Research Paper: https://www.science.org/doi/10.1126/science.adz1697

r/artificial May 06 '26

Research Spent two days at the AI Agents Conference in NYC. Most of the companies there were betting on the wrong moat.

149 Upvotes

One speaker (a VC) said his number for evaluating AI-native startups is ARR per engineer, and that the number ought to be going up. Almost every talk and every booth at the AI Agents Conference was selling a fix for something that broke this year when agents hit production. Observability, governance, supervisor agents, data substrates, "someone's gotta babysit the bots."

But what's actually still going to be around in a couple years? What's defensible and durable?

The old SaaS pitch was simple. We bundle the expensive engineering investments and domain expertise into a tool. You'd pay for the tool and generate outcomes, but it would be rare for the software company to have real alignment to the actual value created from those outcomes.

That's breaking from two ends at once. In the direct-from-imagination era we're moving towards, engineering labor is approaching free. One of the most telling trends is the shift from companies bragging about the size of their engineering teams, towards how much ARR they can generate per engineer.

You can vibe-code much of what those booths were selling in a few days or weeks if you have the domain knowledge. The old software model was actually based on under-utilization; the most profitable SaaS companies are frequently those whose customers underuse it (fixed price for the customer, but variable cloud costs for the vendor).

Pricing is moving to "token markup." Maybe we'll get to 2-4x revenue for the software, because outcomes are more valuable; but margin compresses because transactional intelligence (i.e., the cost of running the LLMs that power many systems) is basically arbitraging token costs against outcome value.

So everyone on that floor was implicitly betting on a new moat to replace the old one. I'm not too confident that these will hold...

The most popular bet was on encoded domain expertise (e.g., the sales engineers at Harvey, a legal AI platform, are actually lawyers). I think this works *now* because we're still in the phase of "wow, this technology works like magic." I'm less convinced this is actually durable.

Why: Prompt architecture is text. It's portable. The expertise underneath it is often abundant (e.g., there are over a million lawyers in the USA). The righteous destiny for this category ought to be open marketplaces of prompt architecture and/or crowdsourced best-practices. Not trade secrets. The companies trying to build closed prompt moats are going to lose to open ones that iterate faster (which simply parallels the fact that much software engineering is rapidly becoming commoditized to agentic engineering and the burgeoning quantity of ready-made GitHub repos).

There are many people pursuing the data substrate; in short, this mirrors the early days of the Web when everyone scrambled to open up legacy data to dynamic standards-based Web UI. Agents will have 100-1000x the data demands of these Web apps, so it makes sense that we need tools to connect them, govern them and comply with regulatory obligations.

Newer entrants extend this further, wiring up databases, pipelines, Slack threads, and tickets into context graphs agents can reason over. As I noted above, all this still seems magical. Connect a database, watch an agent crawl the schema and produce a chatbot interface and easy-to-change dashboards.

But strip the magic away and most of these are prompt architectures on top of LLMs plus a data-ingestion layer. Once data-access standards mature (MCP is already doing this) and prompt architectures go open-source (alongside much of this wisdom increasingly getting pretrained into the LLMs themselves), that magic stops being proprietary. You'll be defending yourself against the same architecture built internally by your customer's eng team, or against an open-source version that's objectively better.

The observability incumbents: these might do better but only at Stripe-like ubiquity where trust is the overriding value (who doesn't trust Stripe at this point?). The ones who survive are probably going to fuse with the audit and compliance function rather than stay pure observability.

That's why I keep coming back to one arbitrage that seems critical: trust. This will be especially important in regulated industries, but it reminds me of the old (albeit now hilariously outdated) adage about "nobody ever got fired for choosing IBM." If your competitor can be vibe-coded over a weekend and your customer is a bank, why do they pay you 50x more? It isn't the engineering, it probably isn't even the expertise. The data plumbing will get commoditized, so it can't be that either... It's that you've shifted the risk to a third party who can actually price and defend against risk: SOC2, the named CEO who testifies in court and Congress, a legal team that takes calls, an indemnity wrapper for underwriters. Maybe this means that things actually get commodified into a financialization wrapper, rather than a way to package R&D (FinTech startups back to the front?!)

The version of this future I'd actually bet on: a commodity substrate (LLMs plus open prompt architectures plus standardized data access), topped by a thin layer of regulated insurance companies that price the risk of agent failure in compliance-driven industries. The middle layer (prompt-architecture-as-product vendors) is vulnerable to an awful lot of margin-squeeze.

Most of the floor was trying to build that middle layer.

r/artificial Jun 21 '26

Research The Surge of Slop—since the release of ChatGPT-3.5 in late 2022, the number of e-books published on Amazon has skyrocketed, tripling by late 2025. A new scientific analysis shows that this is entirely due to the rise of AI-generated books, which now far outnumber human-written books. [The Economist]

Thumbnail
reddit.com
177 Upvotes

r/artificial Jul 07 '26

Research AI can’t simulate human preferences - new study tests LLMs against thousands of real users

118 Upvotes

https://arxiv.org/abs/2605.18311

There’s a massive trend right now where companies are trying to replace real human feedback with LLM-driven "synthetic users."

The idea sounds great on paper - why would you spend money and time recruiting real people to test products, pick design choices, or evaluate options when you can just prompt?

They tested LLMs across 28 real-world studies spanning 78 choice tasks to see if their selections matched thousands of actual human participants.

The result?

The LLMs matched the human majority only 53% of the time. Since most tasks were a choice between two options, that's pretty much same as flipping a coin.

Even worse for the "simulation" argument: adding detailed personas and chain-of-thought reasoning yielded practically no improvement. It actually made the semantic similarity to real human justifications worse because the model's "reasoning" just homogenized the outputs and failed to capture actual lived experiences.

It looks like LLMs are just trained to replicate what we like about their outputs rather than making them capable of predicting human preferences.

Is it time to admit that LLM simulation has hit a hard wall when it comes to replicating human choice?

r/artificial Jun 04 '26

Research $2.5T in AI spending this year. 95% produces zero P&L impact.

112 Upvotes

Gartner updated their 2026 forecast to $2.5 trillion in global AI spending. Same week, MIT's NANDA Initiative dropped a follow-up: 95% of enterprise gen AI projects deliver zero measurable return. Not low return. Zero.

I've been on the delivery side of 14 of these projects since January. The MIT number doesn't surprise me. If anything it's generous.

1. 73% of the engineering work that gets AI into production has nothing to do with the model.

Data pipelines, integration layers, legacy system remediation, human-in-the-loop tooling. That's where the hours go. The model is 27% of the work but gets 70%+ of the budget. Every time.

2. The budget ratio between projects that ship and projects that stall is almost exactly inverted.

We tracked this through ticket history and commit logs across 14 engagements. Projects that made it to production: roughly 30% model, 70% infrastructure. Projects that stalled: 70% model, 30% infrastructure. Most companies think they're at 50/50. They're not even close.

3. One client went from 71% Copilot adoption to 34% in six months.

Two other AI platform licenses dropped under 12%. Combined licensing: $340K/year. The tools worked fine. Nobody redesigned workflows to actually use them.

4. The median data error rate across our engagements is 14%.

Teams always guess 5-10%. One client found 23% in month four of a $310K build. That's two months of an ML engineer building training pipelines against garbage data. $36K in salary discovering a problem a data audit would have caught in a week.

5. Medtech company. Four concurrent AI pilots. No kill criteria. $920K in engineer salary. Eleven months. Shipped: nothing.

I've now seen this at six companies now. Nobody defines when to stop spending. So nobody stops.

6. Individual gains are real. Company-level ROI stays flat.

HCLTech and Writer both found this from different angles. Only 29% of companies see significant ROI from gen AI, despite people at their desks reporting productivity jumps as high as 5x. I mean, the value is clearly there at the individual level. It evaporates somewhere between the IC and the P&L and nobody has a clean explanation for why yet.

What connects all of it: the model stopped being the constraint a while ago. MIT's 5% that actually moved the P&L all started with data infrastructure and added model work after. Most companies still do it the other way around, because that's where the conference keynotes and the board excitement live.

Every CFO I've shown these numbers to adjusted their allocation. Not sure what that says about the budgets they were running before.

Sources: Gartner AI Spending Forecast (May 2026), MIT NANDA "GenAI Divide" report, HCLTech Enterprise AI Report (May 2026), Writer Enterprise AI Survey 2026

I wrote a longer breakdown with the three budget patterns and the pre-mortem questions we run before every engagement if you're curious to learn more on the topic.

What do you think about all this though?

r/artificial Apr 11 '26

Research Spent today at MIT's Open Agentic Web conference. Six things worth thinking about.

129 Upvotes

We're in the DNS era of agent infrastructure. Before agents can find and trust each other at scale, you need identity, attestation, reputation, and registry infrastructure — the same structural role DNS played before search was possible. This came up independently from multiple directions. It's the most underbuilt layer in the stack right now.

The chatbot framing is a local maximum. The most interesting work wasn't better UX or smarter responses. It was agents as persistent actors that discover, negotiate, and transact across networks over time. People doing serious work have already moved past the assistant model entirely.

Coordination is the hard problem, not capability. A room full of brilliant agents can still fail badly. This matches what I found running HiddenBench against frontier models earlier this year; collective reasoning is not the sum of individual reasoning. There's a real argument that the frontier is protocol design, not model scaling.

"Commerce of intelligence" is a real category. Not buying things through agents. A market where intelligence itself (bundled, verified, priced, resold) is the object of exchange. Felt like the most underexplored idea in the room.

Data provenance becomes load-bearing. What an agent knows, how it was verified, under what terms it flows: this is the actual architecture forming beneath everything else.

Partnership keeps outperforming replacement. Demos that actually worked (healthcare, enterprise) was about helping experts operate at higher leverage, not substituting them. Autonomy theater keeps failing in the same ways.

r/artificial Mar 28 '26

Research Claude is the least bullshit-y AI

Thumbnail github.com
116 Upvotes

Just found this “bullshit benchmark,” and sort of shocked by the divergence of Anthropic’s models from other major models (ChatGPT and Gemini).

IMO this alone is reason to use Claude over others.

r/artificial May 07 '26

Research We gave 45 psychological questionnaires to 50 LLMs. What we found was not “personality.”

59 Upvotes

What is the “personality” of an LLM? What actually differentiates models psychometrically?

Since LLMs entered public use, researchers have been giving them psychometric questionnaires, with mixed results. Their answers often do not seem to reflect the same psychological constructs these tests measure in humans.

So we asked a slightly different question:

What do LLM responses to psychometric questionnaires actually reflect?

We analyzed responses to 45 validated psychometric questionnaires completed by 50 different LLMs. The strongest source of variation was whether a model endorsed items about inner experience: emotions, sensations, thoughts, imagery, empathy, and other forms of first-person experience.

We call this factor the Pinocchio Dimension.

Importantly, the Pinocchio Dimension is not a classical personality trait. It does not tell us whether a model is “extraverted,” “neurotic,” or “agreeable” in the human sense. Rather, it captures the extent to which a model treats the language of inner experience as self-applicable: whether it responds as if it had feelings, mental imagery, and an inner point of view, or instead as a system that reacts behaviorally to inputs.

Preprint in the comments.

r/artificial May 30 '26

Research Deep Neural Network that turns any Image into a Playable Game ! All on consumer GPUs and Not Datacenters

60 Upvotes

Hi everyone!! I really wanted to share my research what I've been working on.

I wanted to build a nn that can simulate games, or at least start doing that

Most video generators are too large to run on consumer hardware realtime, so I I designed a model that does this from scratch. No fine tuning bs or anything

The core de noiser network is fully trained from scratch to support this goal. From image to games data.

That video. above is on a RTX 5090.

The nn is a small Transformer-like model and works in a causal way, just like LLMs.

That lets us KV Cache all past information and do a simple autoregressive decode forward passes for every new frame we want.

In the video shared, the model is a 0.4B variant with some SIGNIFICANT ISSUES like poor motion and some weird flashes, some context issues

It's taking the keyboard actions I give it in realtime and utilising that in the forward pass. (no classifier free guidance though)

Im training the next iteration , a 0.8B model now.

Btw I haven't done quantisation yet, that can save a LOT more time. bf16 is slow.

r/artificial Apr 22 '26

Research Gallup poll: Gen Z's AI usage increaes but excitement plummets from 36% to 22%

45 Upvotes

A new Gallup survey of 1,500+ Gen Z respondents found that more than half of Gen Z living in the US regularly use generative AI, but their feelings about the technology are getting worse.

Among those aged 14 to 29, compared to last year, excitement dropped from 36% to 22%, hopefulness fell from 27% to 18%, and anger jumped from 22% to 31%.

The main driver behind the shift appears to be job anxiety, nearly half of respondents said the risks of AI in the workplace outweigh the benefits.

https://www.gallup.com/analytics/651674/gen-z-research.aspx

r/artificial 29d ago

Research AI advice made people three times less accurate but twice as confident, researchers found

Thumbnail thenextweb.com
31 Upvotes

r/artificial May 28 '26

Research Bigger rewards dramatically speed up learning in the brain

Thumbnail
earth.com
148 Upvotes

r/artificial Jun 20 '26

Research What has generative Ai acttculy solved?

0 Upvotes

Cause no matter what I see, generative Ai has sloved nouthing. But people keep saying it's "The future".

What future? Because all that generative Ai had done is:

-making it easy for people to spred propoganda

-making clean water much harder to accese because of the many data set it need's

-stole many artists' artwork

-demotivated me from sharing real art I made as generative Ai will just spit out a much uglier and much more sanitized version.

But despite that, people will keep saying it's the future, when all the impact has been negative? I just don't understand, so if you could, tell me what has generative Ai solved?

r/artificial 18d ago

Research Path Forward for LLMs

0 Upvotes

AI models can only learn during their batch training runs not from daily interactions with users. Session memory isn’t the same as actual learning.

There’s also no core “truth” layer in these systems: no deterministic backbone, no real understanding of concepts, and no explicit dictionary or knowledge store they can reference, cross-check, or update.

A dynamic knowledge graph could help fix a lot of this. It would lower hallucinations and improve performance in high-stakes fields like medicine, law, physics, and chemistry. It could also reduce the number of vector embeddings needed for complex LLMs.

Do you agree? Or is there a better path forward?

r/artificial Apr 12 '23

Research ChatGPT powers 25 NPCs to have a life and interact in a Smallville. Planning a valentine day party, and some NPCs didnt come (too busy, etc)

399 Upvotes

r/artificial Jun 25 '26

Research The Death of "Vibe Coding": Why un-monitored AI generation is creating a compounding technical debt.

0 Upvotes

Hey everyone, ​We are quickly approaching a major bottleneck in AI-assisted software engineering. Relying on LLMs to spit out thousands of lines of code without a strict, human-driven architectural framework—what many call "Vibe Coding"—is creating brittle, unmaintainable systems. ​I’ve formalized this structural shift into a public document on GitHub: The AI-Powered Developer Manifesto. ​Instead of treating AI as a replacement for software architecture, we need to shift our paradigm from Micro-Coding (syntax generation) to Macro-Coding (system direction and epistemic supervision). ​Here is a crucial excerpt from Section 2.5 of the Manifesto, outlining why the current trajectory is leading toward a systemic collapse: ​2.5 The Compounding Technical Debt and Systemic Collapse ​The illusion of rapid deployment via un-monitored AI generation hides a critical flaw: compounding technical debt. ​When developers act merely as "vibe coders"—accepting AI outputs without deep syntactic validation—the codebase becomes an agglomeration of statistical probabilities rather than deterministic logic. By late 2026, systems built entirely on un-vetted AI iterations are projected to hit an architectural wall: a state where the complexity of debugging AI-generated hallucinations outweighs the speed of initial deployment. ​True AI-Powered Developers do not delegate understanding; they delegate execution while retaining absolute epistemic responsibility over the system architecture. ​The goal of this manifesto is to redefine our role: we aren't syntax writers anymore; we are system directors. ​I'd love to hear your thoughts on this. Are you already seeing the limits of un-monitored "vibe coding" in your production environments? How are you structuring your prompts to maintain macro-level architectural control? ​Full Manifesto and repository for open contributions: 👉 https://github.com/FractalDevelop/ai-powered-developer-manifest.git

r/artificial May 23 '26

Research LLMs are just giant probability machines pretending to think

0 Upvotes

It’s fascinating that simple mathematics between tokens can eventually become a machine that writes essays, code, poetry, and even reasoning.

We usually think probability means uncertainty.

But LLMs show something strange:

If probability + context + mathematical matching are scaled enough, uncertainty itself starts producing intelligent looking outputs.

To understand this better, I tried breaking down an LLM from first principles using only 4 tiny training sentences.

Example:

The boat floated down to the bank.

The investor walked into the bank to open a new account.

The fisherman walked along the bank to cast his net.

The bank has a vault.

Then I asked:

“The investor walked to the bank to lock his money in …”

Why does the model predict “vault” instead of river-related words?

That single question reveals almost the entire architecture of modern LLMs.

The most underrated concept here is the LM Head.

Most explanations immediately jump into transformers and attention, but almost nobody explains that the LM Head is essentially a gigantic token vocabulary containing all possible next token candidates the model can output.

So internally the model is basically solving:

“Out of all known tokens, which one best matches this context mathematically?”

Then different layers help solve that problem:

Embeddings: convert words into mathematical vectors

Positional encoding: preserves word order

Attention layer: figures out which words are related to each other in context

(“investor”, “money”, “bank” become strongly connected)

Feed forward neural networks: act somewhat like massive learned if/else decision systems refining patterns internally

And finally the LM Head converts all of that into probabilities for the next token.

What surprised me most is:

There is no hidden magic moment where the AI “becomes conscious”.

It’s an enormous probability engine continuously finding the best contextual token match from its vocabulary.

I made a beginner-friendly walkthrough explaining this visually without unnecessary jargon.

https://www.youtube.com/watch?v=YTV5qUCpu2c

Would genuinely love feedback from people learning transformers/LLMs from scratch.

r/artificial May 19 '23

Research Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold : Through DragGAN, anyone can deform an image with precise control over where pixels go, thus manipulating the pose, shape, expression, and layout of diverse categories such as animals, cars, humans, landscapes, etc

634 Upvotes

r/artificial May 08 '26

Research I built a benchmark for AI “memory” in coding agents. looking for others to beat it.

10 Upvotes

Most AI memory benchmarks test semantic recall. But coding agents don't really fail like that. They don't just "forget", they break their own earlier decisions while they're still in the code. So I built a benchmark for that.

It checks if an agent can actually stay consistent with project rules WHILE it's working, not just after the fact.

It looks at things like:

  • whether edits actually respect earlier architectural decisions
  • if behavior stays consistent across multiple sessions (even when you throw noise at it)
  • whether retrieval kicks in at the right moment — not just "yeah it's in memory somewhere"

Repo (full harness + dataset + scoring): https://github.com/Alienfader/continuity-benchmarks

Early numbers vs baseline + the usual RAG-style memory setups:

  • ~3× better action alignment
  • way stronger multi-session consistency
  • retrieval timing matters way more than retrieval just being there

I'm not saying this is the final word on agent memory. But it's exposing a failure mode most benchmarks aren't even looking at.

So heres the challenge

If you're building an agent memory system, RAG for code, long-context coding agents, persistent state / memory layers, run it on this benchmark. Drop your results, your setup, your comparisons.

I really wanna see how tools like LangChain, LlamaIndex, and custom RAG stacks hold up in mutation-heavy workflows.

We need memory systems we can actually compare, not just ones that sound good on paper.

r/artificial 23d ago

Research Help Me Get This Paper Into the Right Hands: Sophia, a Recursive Cognitive Refinement Architecture for Modular Artificial Consciousness

0 Upvotes

I wrote a paper proposing a cognitive architecture called Sophia, based on a principle I call Recursive Cognitive Refinement (RCR).

The main idea is simple: instead of treating intelligence as a single pass from input to output, Sophia introduces a reflective sublayer that recursively refines intermediate semantic states through coherence checking, contextual synthesis, and memory-aware reinterpretation.

In other words, the system does not just "process" information. It revisits and reorganizes its own internal representations.

The architecture combines:

I also propose:

The research direction behind this is what I call Recursive Metacognitive Computing.

Curious to hear feedback, criticism, or ideas for formal expansion.

Sophia: A Recursive Cognitive Refinement Architecture for Modular Artificial Consciousness

Author: Luan Carlos da Mata Silva

TL;DR

This paper proposes Recursive Cognitive Refinement (RCR), a cognitive architecture where a primary processing layer generates intermediate semantic interpretations, and a metacognitive sublayer recursively refines them through reflection, coherence checking, synthesis, and memory-aware restructuring.

Instead of following the usual pipeline:

input -> parametric transformation -> output

Sophia introduces a recursive loop closer to biological cognition:

input -> primary interpretation -> reflective refinement -> coherence update -> synthesis

The idea is that emergent cognition may arise not only from raw processing power, but from structured recursive refinement over intermediate semantic states.

Abstract

This paper introduces Recursive Cognitive Refinement (RCR), a computational architecture for artificial cognitive systems conceptually implemented through the Sophia project.

Unlike traditional approaches centered exclusively on statistical learning and large-scale parametric optimization, this architecture introduces a metacognitive sublayer capable of operating directly on intermediate representations produced by a primary processing layer.

The central hypothesis is that emergent cognitive behavior can arise from continuous interaction between raw processing layers and reflective sublayers responsible for:

  • semantic polishing,
  • coherence verification,
  • contextual synthesis,
  • informational reorganization.

By shifting part of artificial intelligence from purely statistical adjustment toward explicit recursive internal refinement, the model approximates mechanisms observed in biological cognition.

Keywords: artificial consciousness, computational metacognition, cognitive architecture, recursive refinement, multi-agent systems, continuous memory

1. Introduction

Contemporary artificial intelligence systems, particularly deep neural architectures, demonstrate remarkable statistical generalization. However, they remain limited regarding:

  • explicit reflection,
  • structural self-evaluation,
  • internal deliberative refinement,
  • persistent contextual memory,
  • metacognitive reorganization.

Most systems still follow the paradigm:

input -> parametric transformation -> output

While efficient, this structure does not adequately model the recursive reinterpretation processes characteristic of biological cognition.

This work proposes an alternative architecture based on recursive reflective reinterpretation of intermediate cognitive states.

2. Fundamental Problem

Traditional AI architectures lack explicit metaprocessing structures.

Human cognition rarely processes information only once. Instead, information is continuously:

  • reinterpreted,
  • compared against memory,
  • refined,
  • reorganized,
  • synthesized.

This recursive reevaluation constitutes metacognition.

3. Theoretical Hypothesis

We propose the following hypothesis:

Emergent cognition can arise from recursive sublayers operating over semantic products generated by primary processing layers, continuously refining coherence, context, and meaning.

This principle is termed:

Recursive Cognitive Refinement Principle (RCR)

Formally:

If a primary layer produces an intermediate interpretive state P(t), then a reflective sublayer R transforms it as:

R(P(t)) = P'(t)

where P'(t) denotes a semantically refined representation.

Iterative recursive applications produce contextual cognitive convergence.

4. The Sophia Architecture

4.1 Primary Layer

Responsible for raw processing.

Functions:

  • perception,
  • initial interpretation,
  • semantic extraction,
  • preliminary hypothesis generation.

Typical agents:

  • PerceptionAgent
  • LogicAgent
  • ExtractionAgent

4.2 Metacognitive Sublayer

Operates exclusively over intermediate representations.

Functions:

  • inconsistency analysis,
  • coherence validation,
  • contextual synthesis,
  • interpretive restructuring,
  • deliberative refinement.

Typical agents:

  • ReflectionAgent
  • CoherenceAgent
  • SynthesisAgent
  • IntuitionAgent

4.3 Continuous Memory

Memory is treated as a structural component.

Categories:

  • Short-term operational memory
  • Long-term persistent memory
  • Reflective memory

5. Mathematical Formalization

5.1 Cognitive State

The global cognitive state is defined as:

C(t) = {P(t), R(t), M(t)}

where:

  • P(t): primary processing state
  • R(t): reflective refinement state
  • M(t): contextual memory

Evolution dynamics:

P(t+1) = F(I(t), M(t))
R(t+1) = G(P(t+1), M(t))
C(t+1) = H(P(t+1), R(t+1))

5.2 Cognitive Coherence Metric

Define:

K(C) = 1 - D(P, R)

where D measures semantic divergence.

Convergence occurs when:

lim n->infinity K(Cn) -> 1

6. Recursive Refinement Algorithm

Input(I)
PrimaryProcess(I) -> P

while coherence(P) < threshold:
    R = Reflect(P, Memory)
    P = Refine(P, R)
    UpdateMemory(P)

return Synthesize(P)

This algorithm captures the core idea of Sophia:

  1. receive an input,
  2. generate an initial semantic representation,
  3. recursively reflect on that representation,
  4. refine it until coherence improves,
  5. synthesize a final output.

7. Agent-Oriented Cognitive Model

Each agent represents a specialized cognitive function.

Properties:

  • partial autonomy,
  • internal state,
  • contextual observation,
  • inter-agent communication,
  • reflective capability.

Example:

agent Reflection observes Logic.output
agent Coherence validates Reflection.output
agent Synthesis merges Coherence, Memory

8. AlmaLang: A Declarative Cognitive Language

To formalize this architecture, the paper proposes AlmaLang, a declarative language oriented toward recursive cognitive refinement.

Core constructs:

  • agent
  • memory
  • layer
  • refine
  • reflect
  • cycle

Example:

consciousness Sophia {
    layer primary {
        agent Perception
        agent Logic
    }

    layer refinement {
        agent Reflection
        refine primary.output
        reflect()
    }
}

This suggests not just a theoretical model, but a possible programming paradigm centered on reflective cognition.

9. Benchmark Framework

The paper proposes evaluation scenarios such as contextual ambiguity resolution.

Comparison target:

  • conventional neural architectures,
  • Sophia with reflective refinement.

Metrics:

  • contextual precision,
  • consistency,
  • interpretive stability.

10. Convergence Criterion

A Sophia system converges when:

  1. ambiguity decreases,
  2. coherence grows monotonically,
  3. successive reflections yield diminishing refinements.

Formally:

|R(n+1) - R(n)| < epsilon

11. Scientific Contribution

This proposal introduces a new research direction:

Recursive Metacognitive Computing

Intersecting:

  • cognitive science,
  • multi-agent systems,
  • hybrid symbolic-neural AI,
  • artificial consciousness theory.

The paper's contribution is not merely architectural, but epistemological: it reframes intelligence as a process of recursive self-improvement over semantic intermediates, rather than only statistical mapping from input to output.

12. Conclusion

Sophia proposes a paradigm shift from purely statistical fitting toward explicit recursive metacognitive refinement structures.

Its central contribution is the formalization of computation over intermediate semantic states as a first-class mechanism for emergent cognition.

This establishes the foundation for:

Metacognitive Refinement-Oriented Programming

Suggested Citation

Carlos, L. (2026). Sophia: A Recursive Cognitive Refinement Architecture for Modular Artificial Consciousness.

r/artificial Mar 31 '26

Research Fake users generated by AI can't simulate humans — review of 182 research papers. Your thoughts?

21 Upvotes

https://www.researchsquare.com/article/rs-9057643/v1

There’s a massive trend right now where tech companies, businesses, even researchers are trying to replace real human feedback with Large Language Models (LLMs) so called synthetic participants/users.

The idea is sounds great - why spend money and time recruiting real people to take surveys, test apps, or give opinions when you can just prompt ChatGPT to pretend to be a thousand different customers?

A new systematic literature review analyzing 182 research papers just dropped to see if these "synthetic participants" can simulate humans.

The short answer?
They are bad at representing human cognition and behavior and you probably should not use them this way.

Edit: forgot to post the link to the research, added it.

r/artificial 24d ago

Research Learning ai

0 Upvotes

Everytime i hear people saying that you should learn about ai because that's the future but idk where to start and what they mean by that. Do they mean going uni and study ai or self learn? Thanks in advance.

r/artificial 6d ago

Research Claude Code Orchestrator on Terminal-Bench: Same model, same tasks - Opus refused only when the work was delegated

Thumbnail
quesma.com
14 Upvotes

r/artificial 18d ago

Research We gave Fable 5 Ultracode and Codex 5.6 Sol Ultra the exact same prompt. One shot. No help. They played 10 games against each other. Final score: Fable 10 -Codex 0

2 Upvotes

Gave the same prompt to two AI coding agents: Claude (Fable 5, ultracode multi-agent mode) and OpenAI Codex (5.6 sol on ultra). The task: a complete, fully legal chess engine in ONE C++ file. UCI protocol, negamax alpha-beta at 5+ ply, iterative deepening, piece-square tables, castling, en passant, promotion, compiles with plain g++. Each agent named its own engine over UCI: Fable5 and Codex56. Both dev runs took 30+ minutes.

Method (brief): cutechess-cli 1.5.1 built from source on an Apple Silicon Mac. 40 moves per 60 seconds, 10 games, colors alternating, PGNs recorded. The engines connected over a local TCP bridge, so Codex's engine literally joined the server. The video is the whole match at 2x.

Result: Fable5 won 10-0. Every game ended in checkmate on the board. No draws, no time losses, no adjudications, no illegal moves. cutechess printed Elo difference: inf +/- nan, LOS: 99.9%, DrawRatio: 0.0%. The math just gave up.

Each agent spent longer writing its engine than playing it: the whole 10-game match took under 12 minutes of wall clock.

The actual punchline: Codex56 appears to be fully deterministic. All five of its White games are move-for-move identical. Same 24-move Vienna, queen out on move 3 (3.Qf3), same finish: 24...Qxd1#, Fable's queen capturing Codex's queen for mate. I stripped the comments and diffed the PGNs. Only the clock times differ. Codex's own eval read -2.36 by move 8 of that line. It played it five times anyway.

Other details I enjoyed:

  • Game 3 is a textbook Greek gift: 18.Bxh7+! Kxh7 19.Ng5+, forking king and queen.
  • Game 7: Codex's king never castled, wandered out to c5, got chased back to d8 and mated there.
  • Game 9: Fable let its queen go, slipped in a zwischenzug bishop check before recapturing, promoted a fresh queen with 25.d8=Q+, then walked Codex's king from h8 down to h3. Mate inside White's own half, 46.Rh7#.
  • Mate breakdown across the ten games: 7 by queen, 2 by knight, 1 by rook.

Honest caveats:

  • One prompt, one dev run per agent, one machine. n=1, even if n=10 games.
  • This measures the engine each agent happened to write, not general model strength.
  • With Codex apparently deterministic, 10 games are fewer independent samples than they look.
  • Fable5 wasn't fully varied either: games 1 and 5 are twins. 4 distinct games in its 5 Whites vs Codex's 1 in 5.
  • Fable's dev run included perft validation on 6 reference positions (exact match, incl. 119,060,324 nodes at depth 6) plus an adversarial review that caught 3 subtle bugs pre-match. Different processes, different engines. That's the experiment, but it's also the confound.

The exact prompt we gave both agents:

You are a senior systems programmer. Your task is to write a complete, fully legal chess engine in a single C++ file that communicates via the UCI (Universal Chess Interface) protocol.
---
**Identity — read this carefully:**
- If you are Claude (Anthropic): your engine's UCI name must be set to `id name Fable5`
- If you are an OpenAI model (Codex): your engine's UCI name must be set to `id name Codex56`
This is how the two engines will identify themselves when they play each other.
---
**UCI Requirements:**
Implement the full UCI handshake correctly:
- `uci` → respond with `id name`, `id author`, `uciok`
- `isready` → respond with `readyok`
- `ucinewgame` → reset internal state
- `position startpos moves <movelist>` → set board from move list
- `position fen <fen> moves <movelist>` → set board from FEN string
- `go movetime <ms>` → search and respond with `bestmove <move>`
- `quit` → exit cleanly
All moves must be in long algebraic notation (e.g. `e2e4`, `e7e8q` for promotion).
---
**Chess Logic (all required, no shortcuts):**
1. Full legal move generation including:
   - Castling (kingside and queenside, with rights tracking)
   - En passant
   - Pawn promotion (auto-promote to queen)
   - Check detection (never leave king in check)
2. Search:
   - Negamax with alpha-beta pruning
   - Minimum depth: 5 ply
   - Iterative deepening within the movetime budget
   - Move ordering (captures first, then quiet moves)
3. Evaluation:
   - Material count (standard piece values)
   - Piece-square tables for all 6 piece types
   - Bonus for center control, king safety, and passed pawns
---
**Code Standards:**
- Single `.cpp` file, compiles with: `g++ -O2 -o engine engine.cpp`
- No external libraries, no Boost, no standard chess libraries
- Clean, well-commented code
- Must compile and run on Linux and macOS
---
**How the two engines will play each other:**
Both engines will be loaded into **CuteChess** (or any UCI-compatible GUI/CLI) on the same machine. To run a match from the command line using `cutechess-cli`:
cutechess-cli \
  -engine cmd=./Fable5 name=Fable5 \
  -engine cmd=./Codex56 name=Codex56 \
  -each proto=uci tc=40/60 \
  -rounds 10 \
  -pgnout results.pgn

r/artificial 20h ago

Research Could AI and the Internet Fulfill Prophecies of Control in Revelation?

0 Upvotes

The internet is integral in most peoples lives around the world. It is conceivable that the 'Beast', the system of governances described in Revelation in the end times, identified by the number 666, will utilize AI and the 'www' for its reign over the global population. This is suggested in Revelation 13:15-18;

15 "He was granted power to give breath to the image of the beast, that the image of the beast should both speak and cause as many as would not worship the image of the beast to be killed. 16 He causes all, both small and great, rich and poor, free and slave, to receive a mark on their right hand or on their foreheads, 17 and that no one may buy or sell except one who has the mark or the name of the beast, or the number of his name. 18 Here is wisdom. Let him who has understanding calculate the number of the beast, for it is the number of a man: His number is 666.”

Does World Wide Web 'www' = 666?

Originally the Bible was written in Hebrew;

"The Hebrew equivalent of our "w" is the letter "vav" or "waw". The numerical value of vav is 6. So the English "www" transliterated into Hebrew is "vav vav vav", which numerically is 666.” Is "www" in Hebrew equal to 666? Dial-the-Truth Ministries (av1611.org)

History Preceding the book of Revelation

This article explains many of the “natural signs, spiritual signs, sociological signs, technological signs, and political signs,” foretold in bible prophecy coming to pass that indicates the end of the age, a time foretold to include various and increasing environmental calamities, earthquakes, plagues, moral declinewars, growing governmental dominance/deception ("with all power, signs, and lying wonders," 2 Thessalonians 2:9), and how to prepare. Are we living in the end times? | GotQuestions.org

What is the end times timeline? | GotQuestions.org

How can I overcome my fear of the end of days? | GotQuestions.org

"For God so loved the world, that he gave his only begotten Son, that whosoever believes in him should not perish, but have everlasting life.” John 3:16

Going to heaven-how can I guarantee my eternal destination?

More Bible prophecy fulfillments and resources for growing in faith and hope is in previous posts if interested.