how-ariadne-compares
How Ariadne Compares — Claude & ChatGPT
A capability comparison between Ariadne and the two leading hosted AI platforms: Anthropic Claude
(Claude.ai, Cowork, Claude Code, Projects, Skills, MCP) and OpenAI ChatGPT (GPT-5.x, ChatGPT Work /
agent mode, Projects, Tasks, Connectors, Sora). Two lists: where Ariadne has gaps in core functionality,
and where Ariadne has capabilities the other platforms don't. Part 1 is an honest improvement backlog —
each gap is tracked as a prioritised workstream in docs/platform-assessment-2026-07.md (§4) — and Part 2
is the raw material for marketing.
Snapshot: September 2026 (revision 3 — originally July 2026).
Part 1 — Core-functionality gaps
Things Claude and/or ChatGPT users take for granted that Ariadne lacks or only partially covers. Status: ⬜ open · 🔶 partial · ✅ closed since the July snapshot.
| # | Gap | They have | Ariadne today |
|---|---|---|---|
| 1 ⬜ | Shareable conversation & artifact links | Both: one-click public share links for chats and generated artifacts | No public share mechanism; conversations are private to the account |
| 2 🔶 | Mobile apps | Both: polished iOS/Android apps with voice, camera, push | The web UI is now an installable PWA (manifest, icons, offline service worker). Still no native app and — the bigger gap — no push notifications, so proactive contact and daily digests can't reach a phone |
| 3 ✅ | Project workspaces | Both: "Projects" grouping chats + files + custom instructions in one persistent container (ChatGPT: up to 40 files; Claude: project knowledge) | Closed: project workspaces bind conversations, work, co-edited documents, files, repositories, whiteboards, data connections, test suites, releases and support tickets to one container, with a charter injected into the system prompt of every conversation in the project. Beyond both: an evidence-checked discovery pass, a timeline, and markdown/JSON export |
| 4 🔶 | Model choice & routing | Both: multiple model tiers selectable per chat, automatic routing (fast vs frontier), cost-tiered fallbacks | Models are now configured per agent and per channel with per-model behaviour flags and per-model pricing. Still missing: a per-conversation model picker and automatic cheap-tier routing for background work |
| 5 ✅ | Packaged deep-research mode | Both: one-button multi-source research with inline citations producing a structured report | Closed: Deep Research — planned sub-questions, multi-round search with its own gap analysis, structural citations, an optional plan-review gate, and an explicit contradictions section. Beyond both: grounding anchors (a URL or note you name is read first and cannot be triaged away), a deterministic caution banner when a run could not ground itself, and reports that are editable afterwards rather than only re-runnable |
| 6 🔶 | Finished-deliverable UX | ChatGPT Work ships spreadsheets/slides/web apps as rendered deliverables; Claude Artifacts render interactive output in a side panel | An artifact panel now docks beside every conversation, collecting the documents, rendered HTML, images, code and files a conversation produced, with one click to promote any of them into a versioned co-edited canvas document. Office/PDF tools produce files, whiteboard/visualise render, and MCP Apps render interactive tool UIs in-chat. Remaining: shipping spreadsheets/slides/web apps as first-class rendered deliverables |
| 7 ⬜ | Real-time voice conversation | Both: low-latency full-duplex voice with interruption, emotion, camera input | The voice pipeline gained hardware satellites, a custom wake word and channel mixing, but it remains a serial cascade: latency and full-duplex turn-taking are below the hosted "advanced voice" bar |
| 8 🔶 | Meeting recorder | ChatGPT Record Mode: meeting transcription + summarisation | Ariadne now attends meetings (calendar-driven, provider-agnostic, Teams audio) — arguably beyond Record Mode. Passive recording/summarising of any ad-hoc meeting is the remaining slice |
| 9 ✅ | Ecosystem / store | ChatGPT: GPT store & 60+ first-party connectors; Claude: MCP connector directory | Closed: an MCP server catalogue with trust tiers, official-registry sync, one-click install, OAuth 2.1 with dynamic client registration, and per-conversation server selection. A persona/template gallery would still be nice |
| 10 🔶 | Enterprise trust artefacts | Both: SOC 2 / ISO 27001, DPAs, data-residency options, admin consoles, SCIM provisioning, audit-log export | Self-hosting sidesteps some of this. An operator admin console now exists for billing; SCIM, audit-log export and compliance certification remain open for the hosted offering |
| 11 ✅ | Inline co-editing (canvas) | Both: side-by-side document/code canvas the model edits with tracked changes | Closed: an in-chat canvas panel the agent can dock by itself, with live agent edits, tracked revisions and attribution, and a conflict path that never silently overwrites what you are typing. The same canvas serves Notes and project documents. Suggest-mode and presence/cursors are still to come |
| 12 ✅ | Cost optimisations | Both platforms benefit from prompt caching, batch APIs, cached-input pricing | Closed: provider-aware prompt caching (Anthropic/Azure/DashScope), byte-stable prompt prefixes, cache-preserving abridgement, and cached-input tokens metered and billed. Batch APIs remain unused |
| 13 🔶 | Onboarding & limits UX | Both: guided onboarding, graceful rate-limit messaging, usage meters in-product | Usage meters and real quota enforcement now exist (billing). Guided first-run is still missing |
| 14 ⬜ | Browser extension | Both ship browser extensions/side panels | Desktop computer-use covers some of this, but no lightweight browser extension |
| 15 ⬜ | Image/video generation breadth | ChatGPT: Sora video, high-end image models built in | Image generation exists via tools; no video generation |
| 16 ✅ | Personalisation that models you | Both ship cross-session memory that adapts to the individual | Closed: an evolving model of the person — beliefs with evidence, per-user traits, mannerisms and preferences, behavioural signals, and a nightly synthesis pass. Reviewable and correctable on the About You page. See Part 2 for why this goes further than either competitor |
Not gaps (sometimes assumed to be): scheduled/recurring tasks (cron schedules and time triggers), code execution (Python sandbox with safety review), computer use (local/VNC/browser mediums), persistent memory (arguably deeper than either competitor), file knowledge/RAG (blob indexing + pgvector), custom public bots (personas), MCP support (both client and server side, at runtime), connector catalogue (5,000+ entries with trust tiers and one-click install), project workspaces, deep research, inline canvas co-editing, prompt caching.
Part 2 — Capabilities the other platforms don't have
The marketing list: things Ariadne does that neither Claude nor ChatGPT offers today.
Genuine autonomy, not scheduled prompts
- Goal-driven continuous operation — long-term goals with success criteria, plans, and journals, worked in idle and background lanes by a global scheduler. ChatGPT Tasks and Claude's scheduled sessions re-run prompts; Ariadne pursues outcomes across days and verifies them.
- Evaluator-verified completion — an independent evaluator agent must agree the success criteria are met before a goal closes. Neither competitor has machine-checked "done".
- A perception layer — ten kinds of watcher (RSS feeds, web-page changes, calendar lookahead, upcoming meetings, infrastructure thresholds, support tickets, note and whiteboard edits) emit observations that a triage stage turns into knowledge, notifications, or work items — with per-source rules you control and the ability to re-triage anything after the fact. The system notices things without being asked.
- Proactive contact policy — per-user channels, urgency thresholds, quiet hours, and rate limits govern when the AI may interrupt you; everything else lands in a daily digest of what it did while you were away.
- Enforced autonomy budgets — per-goal daily token budgets, failure backoff, and a full activity feed of every autonomous action with its token cost. Autonomy with a meter and a brake.
A system that improves itself
- Runtime tool creation — the agent writes, compiles (Roslyn), loads, and uses new C# tools during a conversation. Competitors' toolsets are fixed between releases.
- Nightly self-optimisation with a validation gate — skills and knowledge are refined from real usage, and proposed skill edits must beat the old version on held-out evidence before being accepted.
- Memory consolidation — nightly deduplication, decay, and promotion of episodic memory into durable knowledge, so the world model improves rather than accretes.
- Agent-derived goals — agents can form their own (budgeted, reviewable) goals from what they learn about the people they work with.
- Self-reflection mid-turn — the agent can examine its own reasoning while working, catching unverified assumptions and non-convergent paths before they become wrong answers.
- Value-ranked work selection — when several goals are runnable, the scheduler picks by expected value against cost rather than by whatever came first.
A model of you that can be wrong — and knows it
- A falsifiable user model — Claude and ChatGPT remember facts about you. Ariadne holds beliefs with the evidence behind them, a confidence level, and the conditions under which they apply, then argues against itself nightly (thesis → antithesis → synthesis) and tests the result against evidence it held back. Beliefs that stop being true are invalidated with a date rather than quietly overwritten.
- Facets — the different selves you present — the same verified person can behave differently on a Discord server and in a terminal, and the model treats that as a first-class distinction rather than flattening it into one averaged personality.
- Personality, traits and mannerisms, per person — how you like to be addressed, the register you prefer, the habits you display; monitored from real interaction, not filled in on a settings form.
- It asks when it isn't sure — where a belief is genuinely contested, the agent will ask you to settle it — once, at most one belief a night — instead of silently guessing.
- You can read and correct the whole thing — the About You page shows what the system believes about you and why. External connectors may submit evidence, never beliefs: there is deliberately no write-my-own-conclusion API. Off by default, and per agent.
One brain, every channel
- Omnichannel presence with shared memory — web (installable PWA), Windows desktop, CLI, IDE (ACP), Discord servers, Microsoft Teams (including meeting audio), Matrix, email (send and triage inbound), telephone calls via SIP, voice satellites, and IoT physical accessories (buttons, lights). The same agent, same memory, on the phone and in your terminal.
- Your own voice hardware, no cloud in the middle — Ariadne speaks the ESPHome native API directly to Home Assistant Voice PE devices (no Home Assistant required, no reflashing), including browser-based flashing and provisioning over Web Serial and a custom "Ariadne" wake word.
- Embeddable public personas — publish a capability-scoped, rate-limited, origin-locked chat widget of your own agent on any website, no login required for visitors.
- Your instance is an API — and an MCP server — an OpenAI-compatible
/v1/chat/completionsgateway means anything speaking the OpenAI API can use your agents; and Ariadne exposes itself as an MCP server, so Claude or any MCP client can query your memory, notes and tools as a context lake.
Team of agents, unified workspace
- Multi-agent collaboration — named agents with distinct skills and memory converse with each other, delegate with evaluator oversight, and escalate reasoning depth on demand.
- A unified work graph — goals decompose into kanban work items shared between humans and agents; plan steps, journals, and comments live on one board both can see and edit.
- Collaborative surfaces — real-time shared whiteboard (agents draw and rewrite), 3D visualisation workspace, notes with live co-editing, and scheduled topic briefings — all agent-writable.
Runs your operations, not just your chats
- Support ticketing with an agent on the queue — tickets are a first-class object the agent triages and works, and new tickets feed the same perception loop as everything else.
- Infrastructure and deployment — connect Docker hosts and Portainer environments, run remote coding sessions, author and roll deployments, and diagnose live systems with evidence-gated remediation; infrastructure thresholds are themselves watchers, so the agent notices problems before you do.
- Automated testing as a first-class capability — a test registry with provenance, change-targeted runs, and results the agent reads and acts on.
- Metered, billable, multi-tenant out of the box — prepay credit ledger, per-model rating including cached-input tokens, plan allowances, spend caps and Stripe top-ups. Neither competitor lets you run an AI business on their platform; Ariadne is the platform.
Ownership and transparency
- Self-hosted, private, model-agnostic — run the entire platform on your own hardware, keep every conversation in your own PostgreSQL, and point it at any OpenAI-compatible model endpoint. No vendor lock-in on the model or the platform.
- Total token transparency — every token, including all background and autonomous work, is attributed to a user, agent, model and purpose, and is queryable per person and per organisation.
- Guard rails you control — a configurable classifier engine gates agent behaviour, with per-persona capability options for anything public-facing.
- Browser automation with credential isolation — recorded, self-repairing Playwright flows where the model never sees stored credentials.
- The platform develops itself — an Ariadne instance actively works on Ariadne's own codebase; the autonomy features are dogfooded on real software-engineering work daily.
Related reading: Goals & Schedules, Personas & Web Chat Widget, Agents, Account & Organisation.