--- name: capability-exploration description: Explore current products, enterprise systems, models, developer interfaces, and available host tools to discover and test capabilities that improve an AI solution. Use for open-ended technical exploration, platform or model choices, buy/adapt/build decisions, and finding a maintainable way to deliver a specific workflow. --- # Capability exploration Discover what can usefully be done in this environment and develop the best path to the requested outcome. Research, inspection, and experiments should change the solution or settle a consequential choice. Use the user's requirements and existing decisions as a starting point while remaining alert to capabilities that enable a better experience than the brief anticipated. Work at the level of the decision. Selecting a model inside an approved application, extending a licensed product, and designing a new service have different scopes. Preserve chosen platforms unless the task includes reconsidering them or a demonstrated limitation requires discussion. Make a recommendation when the available evidence supports one and complete authorized exploratory work rather than stopping at a comparison plan. ## Start with a task worth trying Choose a representative task or situation that exposes what the solution must accomplish. Recover its input, useful outcome, environment, and consequential variation. A concrete renewal case, disputed invoice, engineering specification, or investigation provides a better basis for exploration than a list of desired AI features. Identify the uncertainty that could most change the design: interpreting a difficult document, preserving record access, assembling information from several systems, keeping a long task alive, creating an editable artifact, performing a business operation, or supporting the needed interaction. Follow the uncertainty in this task; these examples are not a prescribed capability taxonomy. Keep space for invention. A requirement such as “summarize these cases” may describe what the user already believes possible. Discovering that the environment can compare revisions, resolve business entities, or continue work after a new event may suggest a more useful service. Explain the proposed change in terms of the user's outcome and retain the original objective. ## Enter the environment through its useful surfaces Inspect what is available to the running host: connected tools, callable metadata, relevant skills, supplied files, existing applications, repositories, and authorized account or browser surfaces. Use purpose-built interfaces where they expose the required behavior clearly. Read a capability's actual description and input/output semantics before relying on its name. Missing exposure to this host does not establish that the underlying product lacks the capability. Explore the customer's environment as well as the market. An installed system may already own a case lifecycle, document relation, approval, or operator view that would be expensive to recreate. Its useful capability can be visible in different places: - The operator experience shows how work is opened, inspected, changed, and completed. It reveals context available at the moment of use and often exposes product functions absent from a generic API overview. - Administrative and configuration surfaces show enabled features, entitlement, scope, identity, extension settings, and operating constraints. A visible feature in a public demonstration does not establish that the customer's account can use it. - Developer and integration surfaces show accessible objects, events, supported extensions, callable operations, execution behavior, and deployment boundaries. Inspect examples and schemas that touch the task, including their limitations. - Operational material shows how the capability is supported, upgraded, observed, and recovered when something fails. This matters when the solution will depend on it repeatedly. Traverse the surfaces that answer the evolving question. A discovered action may lead to its object schema, which may reveal an existing business relation, which may remove the need for a separate data model. A supported extension may alter the interface design. A durable case function may change where task progress belongs. Follow promising affordances far enough to understand their implications instead of freezing the architecture before exploration. Use the access the user has provided. Discovery can proceed through read-only inspection and reversible local work; account changes, installations, production actions, and paid experiments follow the task's existing authorization. If an inaccessible surface controls the decision, identify the exact question and the least burdensome way to answer it while continuing independent investigation. ## Search outward where a gap or a possibility warrants it Form searches from the work and the unresolved behavior, including the relevant enterprise ecosystem. Search for ways to accomplish the outcome as well as the name of a presumed feature. The useful result might be a native product function, a maintained extension, a different interaction, a model capability, or a combination that was absent from the initial design. Read current primary technical material for promising candidates, then inspect relevant examples, feature limitations, release or migration notes, and the selected deployment surface. Product availability, versions, licensing, retention, quotas, and pricing change; settle the facts that affect this choice at task time. A public claim can justify investigation, while detailed documentation and a representative probe establish what the design may rely on. Track the exact configuration where it matters. A named model exposed through one product may have different tool access, context handling, lifecycle, or regional availability through another. A function can exist in an interactive product but lack the required unattended interface. An inference endpoint, a file store, and a trace service can have different data handling. Follow the proposed flow far enough to see those differences. Investigate alternatives proportionately. Search beyond familiar names when a consequential gap remains or an unfamiliar approach could materially improve the result. Stop widening the search when additional candidates are unlikely to change the choice. The output should make the useful discovery accessible; the user should not have to reconstruct it from a bibliography. ## Put promising capabilities to work Use a small but discriminating piece of the workflow to test the capability. Inspect sample outputs, try relevant controls, call available operations with authorized inputs, or build a local experiment where appropriate. Extend the probe when a result reveals another important question. Reading all the documentation before trying anything can leave the most consequential misunderstanding intact. Design probes around the claim the solution needs. If the benefit comes from comparing revised agreements, a single clean agreement tests too little. If it depends on using current operational state, a static fixture cannot establish freshness. If it depends on a case continuing after an interruption, inspect the retained state and resumption behavior. Use the smallest experiment that meaningfully tests the proposition, even when that experiment requires several connected steps. Observe both the capability and its composition with the workflow: - A connector may return permitted documents while its search operation applies a different scope. Try the relevant retrieval path, not just a successful connection. - A long-running operation may finish correctly while leaving the user unable to recover or amend it. Inspect how progress, cancellation, and a changed input are represented. - A model may extract accurate facts while missing the relationship needed for the decision. Examine the complete task output, such as whether the applicable amendment changes the proposed action. - A generated artifact may render attractively while losing the editability, formulas, citations, or downstream import behavior the user needs. Inspect it in the intended receiving environment when possible. - A tool may acknowledge a request while the business operation remains pending. Establish what its receipt means and how the authoritative result is obtained. Use these examples only where they apply. The purpose is to find what the proposed service can rely on and where its design must adapt. When testing intelligence, first establish a credible capability ceiling with a configuration suited to the difficult work. Then explore cost, latency, and operational fit without discarding necessary quality. A weak initial configuration can create a false conclusion that the task is impossible; an expensive successful demonstration can conceal unacceptable operating economics. Compare task-appropriate configurations on equivalent outcomes and inspect why they differ. Explain a failed probe by layer. Determine whether the problem comes from unsupported capability, account access, tool exposure, input preparation, configuration, task ambiguity, or the capability's observed limit. An authentication error says little about task quality. A retrieval miss may reflect an indexing or permission issue. Try a materially different hypothesis when the evidence supports it, then revise the design rather than repeatedly retrying the same path. ## Assemble coherent choices Describe each serious option as a working arrangement. Identify where the user works, where business state lives, how intelligence obtains context and performs work, what connects it to the enterprise, and who maintains the parts. Separate decisions that can vary independently: product experience, model, inference location, execution environment, orchestration, data, and integration. Combine them according to the task rather than treating one vendor label as the whole architecture. Consider reuse, configuration, supported extension, custom components, and combinations on their merits. A native product may save substantial integration work but limit the desired interaction. A custom experience may be justified by a high-value workflow while using maintained services underneath. An open-weight model may provide useful deployment or adaptation options, but its host, data flows, license, and operational burden still determine fit. A managed service may remove infrastructure work while creating an important dependency or export constraint. None of these properties settles the entire choice. Compare the tradeoffs that affect accepted work. Include task quality, operator review, integration effort, maintenance ownership, latency and capacity, total cost, lifecycle, and exit where consequential. Derive the comparison from the workflow: an overnight preparation job and a live customer conversation value different behavior. A cheap attempt can become an expensive accepted result if it needs retries or expert rework. Investigate the seam that can overturn the recommended arrangement. If the solution relies on a source's permissions, a case engine's human handoff, or a runtime's resumability, probe that seam before elaborating secondary features. State a working assumption when direct verification is unavailable, show what changes if it fails, and recommend the next action that resolves it. Keep useful design moving around that dependency. ## Worked exploration when a seam determines the design Read [the engineering-workbench probe](references/workbench-probe.md) when an apparently available capability must be tested across a source change, an interruption or an access boundary. The illustrative traces show why a successful tool response can support the wrong architecture, and distinguish what a local fixture establishes from what requires a live product and account. ## Leave the user with a decision and a usable discovery Present the recommended arrangement and why it fits this work. Show the discoveries that changed the thinking, relevant observations from probes, the material alternative, and the uncertainty that could reverse the choice. Preserve supporting technical sources and useful experiment artifacts without making a source register the main result. When a promising capability is proven enough for the next step, apply it to the requested output or carry it into solution design. When it is limited, develop the corresponding boundary or alternative. Revisit the recommendation when the workload, operating constraints, or dependent capabilities change. The durable knowledge is how the parts serve this task and what their boundaries imply, not a standing ranking of products.