Guide
Glossary
Plain-English definitions for AIAO, AI, agents, and the agentic economy. If a term shows up in the UI — or in the conversation around it — it should be here. Billing terms (credits, retainers, pass-through, HITL) live on the dedicated credits guide.
Looking for credits & billing?
Credits, retainers, pass-through debits, HITL approvals and their rates now live on the How credits work guide, next to the worked examples.
AIAO platform
Terms specific to the AIAO platform and its scoring model.
- AIAO (AI Agent Optimization)
- AI search and conversion optimization — but full-funnel and full-stack. Full-funnel: Visibility → Recommendation → Referral → Handoff conversion, across every AI agent that touches your category. Full-stack: audit, optimize, measure, monitor, attribute, host MCP, coach, consult — the whole toolset in one platform. The successor discipline to SEO and ASO for the agentic era, spanning websites, mobile apps and MCP servers / APIs.
- ARO (Agentic Readiness Optimization)
- The discipline of making a website, mobile app, or MCP provider agent-ready — technically and editorially — so AI agents can discover it, select it over alternatives, and trust it enough to recommend it to users asking for information, shopping options, or tools/apps to invoke. Covers schema, manifests, structured content, machine-readable claims, intent coverage, freshness, provenance, and the editorial framing that makes an asset citable. ARO is the domain; the ARO Audit is the 60-second read-only entry point that scores ~250 signals into a DST score across Website, App Store / Play, or MCP manifest.
- ARO Audit
- The read-only, 60-second baseline scan inside the ARO domain. Turns ~250 signals into a DST score on any website URL, App Store / Play link, or MCP manifest — no setup, no connections. Every finding maps to the Optimizer pillar that fixes it.
- ARO Index / AIAO Index
- The 0–100 composite score for a brand asset (website, app, or MCP server), weighted across the 12 pillars. In App mode the headline is the Listing DST score; in MCP mode the Tool DST score. Same scale and bands across all three.
- DST (Discovery · Selection · Trust)
- The three things every agent does before recommending you: find you (Discovery), pick you over alternatives (Selection), trust you enough to act (Trust). DST is the 0–100 score that blends them.
- ABF (AI Brand Footprint)
- The Intel report that shows how each model in the agent panel currently represents your brand — what it cites, what it recommends, what it gets wrong, and what it ignores. The diagnostic counterpart to ARO.
- Asset kind
- Which surface ARO is auditing — Website, Mobile App, or MCP server / API. Picked at the start of every audit so AIAO only scores factors relevant to that surface (no MCP factors on a brand that doesn't ship an MCP, no App Intents on a pure website).
- Agent panel
- The fixed set of 8 frontier agents AIAO queries every cycle: ChatGPT, Claude, Gemini, Perplexity, Copilot, Grok, Meta AI, Apple Intelligence. Same panel across web, app and MCP modes — only the prompt class flips (URL-naming, app-naming, or tool-calling). Every factor in the catalog is mapped to which of these 8 agents it most affects, so a per-agent gap (e.g. missing Applebot-Extended for Apple Intelligence, missing Copilot Studio connector, missing Grok X App Card) is auditable individually.
- Audit cycle
- One full pass of all prompts across the agent panel. Frequency depends on subscription tier.
- Confidence dot / confidence ladder
- Solid / half / hollow indicator showing how complete the sample for a pillar was this cycle. The ladder goes coarse → partial → refined as more cycles complete.
- Mindshare
- AIAO's name for the rank-weighted share of category prompts in which your brand (website, app, or tool) appears in the agent's answer.
- Pillar
- One of the twelve optimization surfaces: AEO, GEO, AAIO, AXO, DAO, PULSE, RIVAL, PROTO, BUG, SIGNAL, MCP, SOCIAL.
- Push-to-Live
- The Optimizer action that ships a recipe directly to production. Instant edge / CMS deploys in Web mode; staged App Store Connect / Play Console drafts and repo PRs in App mode; manifest, schema and registry updates in MCP mode.
- Recipe
- A pre-built, agent-specific fix inside the Optimizer. Each pillar exposes six recipes, one per model family.
- Revenue-at-stake
- The estimated revenue band tied to your current mindshare and agentic visits — the dollar version of your DST score.
- Penalised-by
- A badge marking which model(s) currently down-rank or omit you, with the failing signal attached.
- Citation drift
- How much what models say about you has changed cycle-over-cycle — wording, facts, sentiment, source URLs.
- Hallucination rate
- Share of agent answers about your brand that contain at least one factually wrong or unverifiable claim, measured per model.
- Persona playbook
- A pre-built audience-and-intent profile used to seed prompts in Intel and Optimizer, so the agent panel is queried the way your real users would query it.
- /.well-known
- A standardized URL prefix on your domain (e.g. /.well-known/mcp.json) that agents check to discover your capabilities.
The 12 pillars
What each of the twelve optimization surfaces actually measures and ships.
- AEO (Answer Engine Optimization)
- Making your content the answer agents quote — schema, FAQ blocks, clean H-structure, citeable facts, llms.txt. Applies to websites and to the long-form content surrounding apps and MCP servers.
- GEO (Generative Engine Optimization)
- Showing up inside generative search surfaces (ChatGPT Search, Perplexity, Google AI Overviews, Copilot) — entities, source authority, freshness.
- AAIO (Agentic Artificial Intelligence Optimization)
- The recommendation pillar — being NAMED and CHOSEN by an agent when it recommends something to a human. Same job on every media type: website, mobile app, or store listing. Signals: entity clarity, category-prompt mindshare, listing/page copy that matches intent, structured facts, ratings and editorial coverage. AAIO is NOT App Intents, tool-calling, or MCP wiring — those are the invocation job and live under the MCP / PROTO / AXO pillars. An app can win AAIO with zero App Intents and vice versa; the two are orthogonal.
- AXO (Agent Experience Optimization)
- Trust and clarity signals agents weight when choosing between options — verified identity, transparent pricing, honest disclosures, low-friction handoff.
- DAO (Dynamic Agent Optimization)
- Detecting which AI agent (ChatGPT, Claude, Gemini, Perplexity, Copilot, Grok, Meta AI or Apple Intelligence) is hitting your app, website, or MCP — by user-agent, IP range, or signed agent header — and serving a response tuned to how that specific model parses, ranks, and cites content. Same source of truth for humans; different rendering for each agent: condensed JSON-LD for token-efficient models, expanded prose + citations for citation-hungry models, tool-call shortcuts for agents that prefer to act over read. The opposite of one-size-fits-all SSR.
- PULSE
- Continuous drift monitoring across the agent panel — when a model starts citing the wrong fact, recommending a competitor, or stops citing you at all, PULSE catches it.
- RIVAL
- Competitive intel pillar — who agents recommend instead of you, on which prompts, with which signals, and what it would take to flip the answer.
- PROTO
- Protocol-level integrations agents discover: MCP server, /.well-known endpoints, OpenAPI, sitemap, robots, llms.txt — the machine handshake layer.
- BUG (Bug Hunter)
- The pillar that surfaces hallucinations, broken citations, wrong prices, mis-attributed quotes and other model-side errors about your brand — and the fix-once recipes to clear them.
- SIGNAL
- Off-platform reputation pillar — reviews, ratings, press, forum mentions, third-party datasets, and presence in region-dominant search engines (Yandex, Baidu, Naver, Yahoo! Japan, Seznam, Coc Coc, Qwant) — the signals agents read about you when they don't read you.
- MCP (Agent-tool Server)
- The pillar that exposes your product's tools to agents via the Model Context Protocol — generates the manifest at /.well-known/mcp.json, tool schemas, auth scopes and rate hints, and pushes them to your connected MCP server. Requires an MCP-capable endpoint; AIAO doesn't host the server itself.
- SOCIAL (Social Grounding)
- Social-graph presence agents read from — X for Grok, Facebook / Instagram / WhatsApp for Meta AI, LinkedIn for B2B. Handle consistency, post cadence and authoritative bios that show up in retrieval.
AI search signals
How LLM-powered search surfaces decide what to retrieve, cite, and recommend.
- AI Search
- Search experiences where an LLM generates the answer directly instead of returning a list of blue links — ChatGPT Search, Perplexity, Google AI Mode.
- AI Overview
- Google's AI-generated summary that appears above the classic results for many queries. A core GEO surface.
- Conversational Search
- Search conducted through multi-turn dialogue rather than single keyword queries — context carries between turns.
- Citation Coverage
- The percentage of prompts in your category where at least one of your URLs is cited in the answer. Discovery-stage signal.
- Citation Share
- Your share of total citations across the prompt set, relative to competitors who also got cited. Selection-stage signal.
- Retrieval Visibility
- How often your content surfaces in the retrieval step that feeds the LLM, regardless of whether it ends up cited in the final answer.
- Source Authority
- How much weight a retrieval system gives a domain when ranking candidate sources — driven by link graph, citations, freshness and entity signals.
- Entity Authority
- Trust and recognition attached to your brand as an entity (not just a domain) — Wikipedia, Wikidata, knowledge panel, consistent attribution across sources.
- Knowledge Graph Presence
- Whether you are represented as a node in major knowledge graphs (Google, Wikidata, Bing) with stable identifiers agents can resolve to.
- Entity Resolution
- The process of recognising that different references (URL, app name, brand alias, founder) point to the same underlying entity.
- Entity Disambiguation
- Separating similarly-named entities so the agent picks the right one — Apple the company vs apple the fruit, two SaaS apps sharing a name.
- Local search engines
- Region-dominant engines that agents fall back to (or are tuned for) outside the Google/Bing axis — Yandex in Russia, Baidu in China, Naver and Daum in Korea, Yahoo! Japan in Japan, Seznam in Czechia and Slovakia, Coc Coc in Vietnam, Qwant in France. AIAO's SIGNAL and DAO pillars query the locally relevant engines based on the country you're tracking, so visibility in Seoul isn't scored against Google alone.
App mode
Terms that only show up when AIAO is auditing a mobile app rather than a website.
- App Intents (iOS)
- Apple-only. A declarative Swift framework for exposing app actions and entities to Apple's system agents — Siri, Shortcuts, Spotlight, and Apple Intelligence — on iOS, iPadOS, and macOS. This is the **invocation** surface for Apple-ecosystem agents and lives under the MCP pillar (agent execution), not AAIO (agent recommendation). Wiring App Intents does not by itself make your app more likely to be recommended; it makes it callable once the user is already inside an Apple agent surface. Android equivalent = App Actions; web equivalent = MCP servers / OpenAPI actions.
- Android App Actions / Intents
- Android's equivalent surface for declaring app actions agents and the system can invoke.
- AASA (Apple App Site Association)
- JSON file at /.well-known/apple-app-site-association on your domain that authorizes Universal Links to open your app.
- assetlinks.json
- Android Digital Asset Links file at /.well-known/assetlinks.json that authorizes app deep links from the web.
- Universal Links
- iOS deep links that open the installed app instead of Safari when AASA authorizes them.
- App Clips
- Lightweight, on-demand slices of an iOS app surfaced from links, codes, or NFC.
- ATT (App Tracking Transparency)
- Apple's user-consent prompt for cross-app tracking. Honesty here is an AXO and BUG signal.
- Store listing
- The public App Store or Play page agents read — title, subtitle, description, screenshots, ratings, category.
- What's-new
- The release-notes field on the store listing; a freshness signal agents weight.
- Ratings velocity
- The rate of recent ratings on the store listing. A Selection-stage signal.
- App Store Connect API
- Apple's developer API used by AIAO to read listing data and (with write scope) stage Push-to-Live drafts.
- Google Play Developer API
- Google's equivalent for the Play Store.
- Listing DST score
- App-mode counterpart of the AIAO Index. 0–100, same bands and formula, scored on 200+ factors with explicit per-agent coverage for Apple Intelligence, Gemini, ChatGPT, Copilot, Perplexity, Meta AI, Grok and Claude.
- Install-intent prompt
- An agent prompt that signals clear install intent (e.g. 'which app should I install for X'). A distinct prompt class in App-mode audits.
- Handoff CVR
- The install-funnel resolution rate from agent answer to store to install, derived from instrumented Universal Links / asset links.
- App-mode ARO
- The 200+ factor audit AIAO runs against an App Store or Play listing, with one mapped factor per agent in the panel covering App Intents (Apple Intelligence), App Actions (Gemini), ChatGPT Apps, Copilot connectors, Perplexity actions, Meta AI in-app handoff and Grok X App Card metadata.
MCP / API mode
Terms that only show up when AIAO is auditing an MCP server or public API rather than a site or app.
- MCP server
- An HTTP/SSE endpoint that implements the Model Context Protocol so agents can discover and call your tools without scraping your UI.
- /.well-known/mcp.json
- The manifest file at the root of your domain that agents fetch to discover your MCP server, its base URL, auth flow and capabilities.
- tools/list
- The MCP method an agent calls to enumerate your tools, their descriptions and parameter schemas. The quality of these descriptions decides whether the agent picks the right tool.
- Tool DST score
- MCP-mode counterpart of the AIAO Index. 0–100, same bands and formula, scored on 210+ factors covering manifest, schemas, auth, errors, latency, registry presence and per-client compatibility across Claude Desktop, ChatGPT, Cursor, Windsurf, Continue, Cline, Zed, Gemini Code Assist, Copilot Studio and Perplexity Tools.
- Tool-call share
- % of agent turns in your category where the agent invoked one of your tools instead of a rival's. MCP-mode equivalent of Mindshare.
- Registry coverage
- % of public MCP registries and tool catalogs (e.g. Smithery, MCP-Index, Anthropic registry) that list your server with a complete manifest.
- Tool-selection precision
- Share of agent attempts that pick the right tool from tools/list on the first try. Low values flag description ambiguity or naming collisions with lookalike tools.
- Schema fidelity
- Share of tool calls that succeed without parameter coercion or retry. Tracks JSON Schema completeness, required-field clarity and enum vs. free-text discipline.
- Hallucinated-call rate
- Share of agent answers that describe calling your tools without actually calling them, often citing wrong parameters or fictional methods. The MCP-mode trust metric.
- Token footprint
- Average tokens per tool response. Verbose tools get dropped in long agent loops — DAO and AXO push this down.
- MCP-client compatibility
- Whether your server works in the major MCP clients (Claude Desktop, Cursor, ChatGPT, Continue) without per-client patches.
- MCP Client
- The application that connects to an MCP server and issues tool calls on behalf of an agent or user — e.g. Claude Desktop, Cursor, ChatGPT desktop.
- MCP Host
- The runtime that orchestrates MCP sessions, manages auth, and routes the agent's tool calls to the right server.
- MCP Resource
- Readable content (files, records, documents) exposed through an MCP server so agents can fetch context without scraping.
- MCP Prompt
- A reusable, parameterized prompt template advertised by an MCP server so agents can invoke it by name instead of reconstructing it.
- Tool Registry
- A public directory of MCP servers and their manifests (Smithery, MCP-Index, vendor registries). Coverage here is part of your Tool DST score.
- Tool Discovery
- How agents locate tools — manifest fetch at /.well-known/mcp.json, registry lookup, in-context advertisement, or tools/list enumeration on connect.
- Capability Negotiation
- The handshake where the MCP client and server agree on supported features (streaming, resources, prompts, auth flow) before tool calls begin.
- Transport Layer
- How MCP messages move between client and server — typically streamable HTTP with SSE, or stdio for local servers.
- SSE (Server-Sent Events)
- A one-way streaming transport over HTTP commonly used by MCP servers to push tool-call progress and results back to the client.
- MCP Gateway
- A broker that fronts one or more MCP servers — adds auth, rate limits, observability, and a single endpoint for many tools.
- Tool Authorization
- The permission model controlling which agents and users can call which tools, typically OAuth 2.1 scopes mapped to tool names.
AI foundations
Core concepts behind every LLM-powered product.
- LLM (Large Language Model)
- A neural network trained on massive text corpora to predict the next token. ChatGPT, Claude, Gemini, Llama, and Grok are all LLMs.
- Tokens
- Sub-word chunks LLMs read and emit. Roughly 4 characters of English ≈ 1 token. Pricing, context limits, and latency are all measured in tokens.
- Context Windows
- The maximum number of tokens an LLM can consider in a single call (prompt + response). Anything beyond it is truncated or summarized.
- RAG (Retrieval-Augmented Generation)
- Fetching relevant documents at query time and injecting them into the prompt so the model answers from fresh, citable sources instead of memory alone.
- Fine-Tuning
- Continuing training of a base model on your own examples to shift its behavior or domain knowledge. Heavier and slower than prompting or RAG.
- Prompt Engineering
- Designing the instructions, examples, and structure given to an LLM to reliably get the output you want.
- Deterministic vs Probabilistic Systems
- Deterministic systems return the same output for the same input every time (SQL, traditional code). LLMs are probabilistic: same input can return different outputs. AIAO's job is to keep the probability distribution in your favor.
- Hallucination
- When a model produces a confident statement that is factually wrong or unverifiable. The single biggest trust risk in agent-driven answers.
- Multi-modal
- Models that accept and/or produce more than text — images, audio, video, structured data. GPT-4o, Gemini, and Claude 3 are multi-modal.
- Model Diversity
- Designing for an ecosystem rather than one LLM. Your brand must perform across all 8 agents — ChatGPT, Claude, Gemini, Perplexity, Copilot, Grok, Meta AI and Apple Intelligence — because each indexes the web, and the app and tool surfaces around it, differently.
Retrieval & grounding
How agents fetch the right context before they answer — the plumbing behind every cited fact.
- Embedding
- A numerical vector representation of a chunk of text, image or other content, used so similar meanings end up close together in vector space.
- Embedding Model
- The model that turns content into embeddings. Different agents use different embedding models, which is one reason your retrieval visibility varies across them.
- Vector Database
- Storage optimised for similarity search over embeddings — Pinecone, Weaviate, pgvector, Turbopuffer.
- Similarity Search
- Retrieving content whose embedding is closest to the query embedding, by cosine or dot-product distance.
- Hybrid Search
- Combining vector similarity with keyword (BM25) retrieval and re-ranking — what most production agents actually use.
- Dense Retrieval
- Retrieval driven by embeddings and similarity search — captures meaning but can miss exact terms.
- Sparse Retrieval
- Keyword-based retrieval (BM25, TF-IDF) — strong on exact matches and rare terms, blind to paraphrase.
- Re-ranking
- A second pass that re-scores the top retrieval candidates with a stronger model before they're handed to the LLM.
- Query Expansion
- Generating alternative phrasings of a query (synonyms, related entities, decomposed sub-queries) to widen recall before retrieval.
- Chunking
- Splitting long content into retrieval-sized pieces so embeddings stay focused and the right passage can be matched without dragging in noise.
- Context Engineering
- Designing the full bundle of information handed to the model — system prompt, retrieved chunks, tool schemas, memory — to maximise the chance of a correct answer.
- Grounding
- Anchoring an answer to specific external sources so claims can be cited and verified, instead of relying on the model's parametric memory.
Agents & agentic systems
How agents reason, act, and coordinate.
- Agents
- Software that uses an LLM as a reasoning core to take actions in the world — call APIs, browse, write files, transact — toward a goal.
- Agentic Systems
- Systems composed of one or more agents that perceive, plan, act, and adapt over multiple steps with limited human supervision.
- Tool Use
- An agent's ability to invoke external capabilities (search, code execution, your API) instead of answering from internal knowledge alone.
- Function Calling
- The interface (usually JSON schemas) by which an LLM tells the host runtime which tool to invoke and with what arguments.
- State Management
- How an agent tracks the current task, intermediate results, user inputs, and tool outputs across many turns of reasoning.
- Plan and Decompose
- The pattern of breaking a high-level goal into ordered sub-tasks before execution — the backbone of any non-trivial agent.
- Multi-Agent Systems
- Architectures where specialized agents (planner, researcher, executor, critic) collaborate or negotiate to solve a problem.
- Long-term Memory
- Persistent storage an agent reads from and writes to across sessions — usually a vector store, knowledge graph, or hybrid database.
- Agentic UX
- Interface patterns built for agents as the primary user: structured responses, deterministic actions, machine-readable status, transparent costs.
- Multi-agent Governance
- The policies, guardrails, and audit hooks that constrain how multiple agents communicate, escalate, and commit irreversible actions.
Agent optimization metrics
The visibility, ranking, attribution and revenue measures AIAO reports on across the agent panel.
- Agent Visibility
- The likelihood your brand (site, app, or tool) appears anywhere in an agent's answer — cited, recommended, or named in passing.
- Agent Ranking
- Your position within the agent's recommendation list when it returns more than one option.
- Agent Share of Voice
- Your share of total recommendations won across the prompt set, relative to the competitors the agent also surfaces.
- Agent Discoverability
- How easy it is for an agent to find you in the first place — manifest presence, registry coverage, sitemap, structured data.
- Agent Eligibility
- Whether your brand even qualifies for recommendation under the agent's filters (region, pricing model, integration support, safety policy).
- Agent Trust Score
- A composite measure of how much an agent currently trusts your brand — feeds Selection and the willingness to act without asking the user.
- Agent Conversion Funnel
- Discovery → Selection → Action → Conversion. The four stages every agent walks before you earn revenue from the answer.
- Agent Attribution
- Connecting downstream outcomes (install, signup, purchase) back to the agent interaction that started them, usually via instrumented links and referrer headers.
- Recommendation Rate
- How often, across the prompt set, an agent actively recommends you rather than merely mentioning or listing you.
- Recommendation Quality
- How relevant and accurate the recommendation is — right use case, right tier, no hallucinated features.
- Agent Influence
- The measurable impact an agent's recommendation has on the user's decision, separate from raw recommendation volume.
- Agent Traffic
- Visits and sessions originating from AI surfaces — answer-engine clicks, deep-links from chat, tool-call referrers.
- Agent Referral
- A single click or action initiated on your behalf by an agent (with an identifiable agent referrer).
- Agent Conversion
- A successful downstream action — signup, install, purchase, booking — attributable to an agent interaction.
- Agent Revenue
- Revenue attributable to agent-driven sessions and tool calls. The dollar version of mindshare.
Data, interfaces & protocols
The machine-readable surface area agents consume.
- Structured Data
- Machine-readable markup (JSON, microdata, RDFa) that describes the meaning of content so agents can parse it without guessing.
- JSON-LD
- JSON for Linking Data — the recommended way to embed Schema.org structured data in a page via a single <script> tag.
- Schema.org
- The shared vocabulary (Product, Organization, FAQPage, HowTo, SoftwareApplication…) that LLMs and search engines use to interpret pages.
- MCP (Model Context Protocol)
- The emerging open standard for letting agents discover and call your app's tools, data, and resources — typically advertised at /.well-known/mcp.json.
- API Design
- How your endpoints are named, versioned, paginated, and documented. Good API design is what makes an LLM-driven agent able to use you correctly on the first try.
- Token Efficiency
- The art of conveying maximum useful information per token — shorter schemas, tight error messages, no boilerplate — so agents pick you over verbose competitors.
- Liquid Frontend
- A presentation layer that renders differently for humans and agents from the same source of truth, so structured facts stay aligned across both surfaces.
Agentic economy
How agents move money, traffic, and decisions.
- Agent Commerce
- Transactions initiated, negotiated, or completed by agents on behalf of a user — booking, buying, switching providers — usually via MCP or structured APIs.
- Growth Loops
- Self-reinforcing cycles where one citation, integration, or fact increases the odds of the next. The agentic equivalent of SEO link flywheels.
- ROI Modeling
- Tying agent visibility (mindshare, citations, agent-driven sessions) to revenue so optimization decisions can be ranked by expected return.
Agent commerce
The mechanics of letting agents transact on the user's behalf — checkout, identity, authorization, and agent-to-agent trades.
- Agent Checkout
- A purchase completed end-to-end by an agent on the user's behalf, often via a tool call against your commerce API rather than the human-facing checkout.
- Agent Booking
- Reservations (travel, dining, services) performed by an agent — usually through MCP or a structured booking API.
- Delegated Transactions
- Any transaction an agent executes on the user's behalf under a pre-agreed scope of authority.
- Purchase Authorization
- The user's explicit approval — per transaction, per merchant, or up to a spending cap — that an agent must obtain before completing a paid action.
- Agent Wallet
- A payment identity attached to the agent rather than a specific session — tokenized cards, stored credentials, or a per-agent payment account.
- Agent Identity
- A persistent identifier representing an agent (not the underlying user) so vendors can recognise, rate-limit, and trust it across sessions.
- Agent-to-Agent Commerce
- Transactions negotiated and executed between two autonomous agents — e.g. a buyer agent and a supplier agent — without a human in the loop.
- Agent Marketplace
- A directory where agents (or their capabilities) are listed, discovered, and sometimes priced for use by other agents or hosts.
- Intent Commerce
- Transactions initiated from an expressed user intent ("book me a flight to Lisbon under €300") that the agent then resolves into a concrete purchase.
- Transaction Confidence
- The agent's estimated probability that the transaction it is about to execute matches the user's actual intent — drives whether it auto-completes or asks.
Trust, safety & governance
Keeping agentic systems honest, secure, and accountable.
- Prompt Injection
- An attack where untrusted content (a page, a review, a tool output) hides instructions that hijack the agent. The XSS of the agent era.
- Adversarial Testing
- Deliberately probing your app with malicious or edge-case prompts to surface failure modes before real agents or attackers find them.
- Privacy-Preserving AI
- Techniques (on-device inference, differential privacy, redaction, encrypted retrieval) that let agents act on user data without leaking it.
- Model Layer Crisis Response
- Your playbook for when a model starts citing wrong facts or stops citing you at all — detect, correct the source of truth, re-seed, and verify across the agent panel.
- Identity and Authentication
- How an agent proves who its user is to your app — OAuth, signed tokens, delegated credentials, MCP auth flows.
- Reputation Monitoring
- Continuously sampling what models say about you (facts, sentiment, comparisons) so drift and hallucinations are caught early.
- Safety Disclosures
- Public statements about what your app does and doesn't do with user data, what it can and can't be used for, and known limitations — both for humans and for model providers.
- Model Cards
- Standardized documentation for an AI model: training data, intended use, performance, known biases, and limitations. The nutrition label of AI.
- Audit Trails
- Immutable, queryable logs of agent decisions and actions — who called what tool, with what arguments, on whose behalf, and what changed.
Hallucination signals
The patterns BUG (Bug Hunter) watches for in agent answers about your brand and code. Surfaced at a high level here; the detailed detection matrix lives in the internal Knowledge.
- Phantom libraries
- The model invents a software package that sounds plausible but doesn't actually exist.
- API drift
- The model mixes up how two different tools or vendors work and blends their commands.
- Version mismatch
- The model uses outdated patterns — or invents future ones — that don't match the current release.
- Undefined helpers
- The model calls a function it never defines, so the code references something that isn't there.
- Attribute / module errors
- The model asks an object or library to do something it can't, or imports something the runtime can't find.
- Deprecated syntax
- The model writes in an older dialect the current language version no longer accepts.
- Fake parameters
- The model passes settings or arguments to a real command that the command doesn't actually accept.
- Circular reasoning
- The model treats its own earlier fabrication as fact and keeps building on it.
- Self-reinforcement
- Once a mistake enters the context, the model defends and extends it instead of correcting.
- Context collapse
- With too much input, the model loses the thread and starts drifting from the actual task.
- Linguistic fluency trap
- The output reads beautifully and confidently, masking that it doesn't actually work.
- Low-probability transitions
- The model fills uncertain gaps with a plausible-sounding guess rather than an accurate answer.
- Docstring fabrication
- The model invents documentation to justify code that doesn't behave the way the docs claim.
- Placeholder stubs
- The model leaves "fill-in-the-blank" comments where the hard logic should be.
- Logical inconsistencies
- The output is grammatically correct code that nonetheless does something nonsensical.
- Variable shadowing & scope leakage
- Names collide or values appear where they shouldn't, producing silent, hard-to-trace bugs.
- Hidden dependencies
- The code only works if something else is installed or configured that was never mentioned.
- Over-simplification
- The model returns a generic answer that ignores the specific edge case the user actually has.
Web, App and MCP ARO factors
The Web, App and MCP audits score 130+, 200+ and 210+ signals respectively across Discovery, Selection and Trust, and every signal is mapped to which of the 8 panel agents it most affects — so a Grok-specific gap, an Apple Intelligence gap or a Copilot Studio gap is auditable on its own. The detailed factor lists are kept off the glossary on purpose — full breakdowns are only published in authorized blog posts.