Appkitekt logo
Appkitekt AIOne agent for AI agent optimization
Request demo

Guide

Methodology

We model how agents decide — not what they'll say tomorrow.

Live accuracy · public dashboard

Rolling 30-day hit rate, citation rate, and drift per scored discovery agent against a holdout brand panel. Refreshes daily.

Open Accuracy Dashboard →

In plain English

AI assistants don't give the same answer to the same question twice. Ask ChatGPT the same thing tomorrow and the wording, sources, and even the brand it recommends can change. Anyone selling you a single "real" number for how visible you are is selling you a screenshot.

What doesn't change every minute is the underlying logic each AI uses to decide who to mention: the files it looks for, the sources it trusts, the signals it weighs. That logic changes when the AI companies ship a new version — not with every user question. So we measure the logic, and use it to forecast your visibility instead of pretending to measure something that keeps moving.

Every score and dollar figure on this platform is a forecast, shown as a range with a middle number. Read the trend and the gap versus competitors — not the single-day figure on its own.

Deterministic about logic, probabilistic about results (the technical version)

AI agents are stochastic — same prompt, different answers across users, sessions, and silent model updates. Appkitekt AI is deterministic about the things that don't change minute-to-minute: documented APIs, agent codebases, manifest formats, retrieval rules, ranking heuristics, tool-selection criteria, and the signals each agent reads when it picks a brand. Those are observable from agent code, vendor documentation, controlled simulations, and the shape of the answers each agent actually produces. They change on a release cadence.

We are probabilistic about any single numeric result: tomorrow's exact rank, the precise citation share on Tuesday, the exact query volume in a given hour. For those we publish a forecast — a central estimate from the geometric mean of the modelled band, with the full range visible — instead of a fake exact number.

What we measure

We catalog and score the signals that AI agents actually use to discover, select, and trust brands — signal selection is observed from agent behavior. Appkitekt AI scores the full chain an agent walks before answering about your app, brand or domain: discoverability, selection, trust, usability, and efficiency. Each of our 12 pillars covers one slice of that chain so you can see exactly where you lose the agent.

Under the hood we evaluate 130+ factors for web, 200+ factors for apps, and 210+ factors for MCP / API vendors across those pillars. Every factor is mapped to which of the 8 panel agents (ChatGPT, Claude, Gemini, Perplexity, Copilot, Grok, Meta AI, Apple Intelligence) it most affects, so a per-agent gap — a missing Applebot-Extended allow, a missing Copilot Studio connector, a missing Grok X App Card — surfaces as its own line item rather than disappearing into an average. The full factor list, weightings, and the recipes that turn findings into fixes are proprietary, but a representative sample is below so you can sanity-check the scope.

How agents work, in one minute

Reasoning & loops. Modern agents run a ReAct loop — reason, act, observe, repeat — often with explicit chain-of-thought, a planner that decomposes the goal into ordered sub-tasks, and an executor that carries them out. Stronger setups add a reflection or critic pass that reviews the work and self-corrects before the agent commits. Appkitekt AI scores you for the surface the agent sees on each pass, not the loop itself.

Memory & context. Agents juggle three kinds of memory: working memory for the current task, episodic memory of past interactions, and semantic memory for persistent facts. Because the context window is finite, the host retrieves only the most relevant slice, compresses older turns into summaries, and chunks long sources so the right passage fits. The cleaner and more retrievable your content is, the more of it survives the squeeze.

Evaluation & ops. On the agent side, providers run evals that track task success rate, tool success rate, latency, cost-per-task and drift over time, with guardrails and tracing in the inference layer. AIAO mirrors the same discipline on your side: every audit is repeatable, every recipe has a verifiable lift signal, and drift in how a model talks about you triggers a PULSE alert rather than a silent regression.

Examples of what we check

Web (sample of 130+): canonical & hreflang integrity, robots.txt allow rules per agent crawler (GPTBot, OAI-SearchBot, ClaudeBot, Google-Extended, Applebot-Extended, PerplexityBot, Bingbot for Copilot, Meta-ExternalAgent), IndexNow + Bing Webmaster submission, per-agent social grounding (X for Grok, Meta surfaces for Meta AI, LinkedIn for Copilot, Apple Business Connect, Google Business Profile), llms.txt directives, sitemap freshness, structured data coverage, JSON-LD validity, OpenGraph & Twitter card completeness, citation density per topic, internal link depth to money pages, HTTPS & HSTS, Core Web Vitals (LCP, INP, CLS), server-side rendering for agent user-agents, semantic HTML landmark usage, heading hierarchy, image alt coverage, content recency, entity disambiguation against Wikidata, review-site coverage (G2, Capterra, Trustpilot), Wikipedia presence, knowledge-panel eligibility signals, pricing page clarity, comparison-page existence, glossary / definition pages, redirect chain length, 404 ratio in crawl sample, cookie-wall blocking of agent fetches — and many, many more.

App (sample of 200+): everything in web mode plus — App Store & Play listing completeness, title/subtitle keyword targeting, screenshot & caption coverage, preview video presence, localized listings per locale, category fit, privacy nutrition label completeness, data-safety form parity across stores, rating volume & recency, review-response rate, version changelog cadence, per-agent app handoff surfaces — App Intents + App Shortcuts donations + EntityQuery (Apple Intelligence), App Actions (Gemini), ChatGPT Apps SDK manifest, Copilot connectors, Perplexity actions, Meta AI in-app handoff (WhatsApp Business / Instagram / Messenger), Grok-readable X App Card — Universal Links via apple-app-site-association, Android App Links via assetlinks.json, deep-link route coverage vs. web sitemap, accessibility traits, Dynamic Type & VoiceOver labels, dark-mode parity, sign-in-with-Apple availability, brand → bundle-ID consistency across stores, curated-list placements (Apple Editorial, Play Editors' Choice), regional chart positions — and many, many more.

MCP / API (sample of 210+): presence in public MCP registries and tool catalogs plus per-agent host enablement (Copilot Studio connector, Gemini Extension, Perplexity Tools, Grok via x.ai console — with Apple Intelligence and Meta AI flagged as no-MCP-host-yet so the matrix stays honest), /.well-known/mcp.json and llms.txt manifest validity, OpenAPI / JSON Schema completeness, tools/list response shape, tool-name disambiguation vs. lookalikes, description quality for agent selection, parameter coverage & required-field clarity, enum vs. free-text discipline, idempotency keys, rate-limit headers, RFC 7807 problem-details error format, auth flow (OAuth 2.1, PKCE, API-key rotation), scope granularity, signed-response support, latency p50/p95 per tool, token-footprint of typical responses, streaming support, cancellation behavior, schema-change deprecation policy, SDK parity across languages, per-client compatibility matrix (Claude Desktop, ChatGPT, Cursor, Windsurf, Continue, Cline, Zed, Gemini Code Assist, Copilot Studio, Perplexity Tools), hallucination rate when agents describe your tools without calling them — and many, many more.

These are illustrative — exact factor names, thresholds, and the pillar each one rolls into are part of the customer deliverable.

Why "real" numbers don't exist — and what we publish instead

AI systems are probabilistic, not deterministic, at the output layer. Large language models generate text by assigning probabilities to possible next tokens and sampling among them. Even with the same prompt, the response can vary depending on sampling settings, system configuration, personalisation, conversation history, and silent model updates.

That's why this platform never quotes you a single "measured" rank or a precise daily query count. Instead we publish a central estimate — the geometric mean of the modelled band — alongside the band itself, and we are explicit that the figure is a forecast conditioned on the logic the agents are running, not a measurement of a stable truth.

That logic is stable enough to model. Agent codebases, documented APIs, manifest formats, retrieval rules, ranking heuristics, and the signals each agent reads when it picks a brand change on a release cadence, not a per-prompt one. We reverse-engineer them from agent code, vendor documentation, controlled simulations, sample answers, and observed tool-call behaviour. Numbers grounded in that logic move when the logic moves — not when a single user gets a slightly different sentence.

The honest unit is therefore a distribution observed under controlled conditions, re-sampled often enough that the trend line means something. Read the trend, the gap, and the relative rank across agents. Treat the single-day exact figure as a forecast, not a measurement.

How we collect signal — the 5-layer stack

Because no single channel can capture a stochastic, personalized surface, AIAO triangulates every finding across five independent layers. A signal only enters your report when enough layers agree, and the confidence score you see is derived from how many of them did.

  1. Real devices. Headed sessions on physical and virtualized devices across iOS, Android, macOS, Windows, and major browsers — what an actual user actually sees, including consent walls, app-clip prompts, and store rendering.
  2. Simulations. Controlled prompt panels held at constant temperature, re-run on a rolling schedule across logged-in, incognito, regional, and tier variants so we can isolate which variable moved the answer.
  3. API investigations. Direct calls to provider, store, registry, and structured-data APIs to ground-truth what the agent's retrieval layer can see about you independent of what it chooses to say.
  4. Agent interrogations. Multi-turn probes of frontier agents in their default consumer configuration — brand, category, comparison, and intent prompts — to measure selection, rank, attribution, and refusal behavior.
  5. Hallucination & anomaly detection. Cross-layer reconciliation that flags fabricated facts, mis-attributed quotes, stale citations, and out-of-distribution swings before they land in your report.

We do not average across the agent panel by default: each model weighs signals differently and your performance will legitimately differ across them — averaging hides exactly the gaps you need to see.

Sampling and scoring

For every audit we generate a prompt set sized to your domain — brand, category, comparison, and intent prompts — held at constant temperature and re-run on a rolling schedule so trend lines stay comparable. Mindshare reflects rate of appearance and rank within the answer, not raw click counts.

App mode flips the prompt class from URL-naming to app-naming and adds store-side signals. The panel, triangulation rule, and scoring approach are identical to web mode.

What we deliberately do not do

  • We do not scrape or republish providers' model weights, system prompts, or training data.
  • We do not claim to know any provider's ranking formula. Every score is observational, not reverse-engineered.
  • We do not synthesize prompts that try to manipulate an agent against its own guidelines.
  • We do not use any competitor product name or proprietary metric label as our own.

How fixes get verified

When you Push-to-Live from the Optimizer, we record the recipe, target pillar, and timestamp. Subsequent scheduled audits re-run the same prompt set and attribute the delta to the push. If a recipe fails to move its target pillar across multiple audit cycles, it is flagged and the pillar moves back into review. App-mode fixes verify on the next App Store / Play indexing cycle once the draft is approved.

Intel metrics — what the dashboards show

Intel is the always-on monitoring surface. It probes the 8-agent panel (ChatGPT, Claude, Gemini, Perplexity, Copilot, Grok, Meta AI, Apple Intelligence) on a rolling schedule and rolls the raw probes up into six metric families that appear identically in both Web and App views:

  • Rankings. For each agent, where your brand sits on the leaderboard it volunteers when asked about your category — position, share of voice, and movement vs. the prior snapshot.
  • Conversion Funnel. The drop-off from agent mention → agent selection → agent handoff (web click, deep link, App Intent, MCP tool call). Shows where the chain breaks per agent.
  • Performance. Latency, refusal rate, citation accuracy, and tool-call success when an agent does engage your surface. Quality of the interaction, not just frequency.
  • Brand Footprint. The corpus of mentions, citations, reviews, and store / web entities the agents can actually retrieve about you — the raw material their answers draw on.
  • Trends. Time-series of the metrics above so a regression after a model update, site change, or app release is visible the next audit cycle rather than the next quarter.
  • Intents. The user questions that surface your brand — coverage across the brand, category, comparison, and problem-led intent classes per agent.

See Reading the numbers for how the Rankings, Funnel, and Performance figures are computed and what good/bad thresholds look like.

Optimizer metrics — what each fix is graded on

Optimizer turns Intel findings into Push-to-Live recipes. Every recipe is graded on four axes before it ships, and re-graded after the next audit cycle to confirm lift:

  • Expected lift. Modelled delta on the target pillar (e.g. +6 pts on Trust, +2 ranks on Gemini leaderboard), derived from comparable fixes in our recipe library.
  • Effort. Implementation cost in hours and which team owns it (web, mobile, content, ops) — so trade-offs are explicit before approval.
  • Confidence. How many of the 5 collection layers (real devices, simulations, API, agent interrogation, anomaly detection) agreed on the underlying finding.
  • Verification window. The number of audit cycles required before a recipe is considered confirmed — short for crawler / manifest changes, longer for store indexing or citation-density work.

The full pillar-by-pillar scoring rubric lives in Scoring; the calculation details for individual metrics live in Calculations.

Want the full breakdown?

The exhaustive factor lists, the exact data sources we cross-reference per pillar, and the recipe library are reserved for customers under NDA. If you're evaluating AIAO for a team or agency, contact us and we'll walk you through the parts relevant to your stack.