--- name: ak-add-capabilities description: > Add capabilities to an existing Agent Kernel project. This skill guides you through adding guardrails, tracing/observability, session persistence, knowledge bases, MCP server, A2A server, AG-UI server, pre/post hooks, multimodal support, conversation thread support, scheduled tasks (deferred and recurring chat execution), the sandbox capability (isolated code execution), and secret resolution (environment first, then AWS SSM Parameter Store or a custom provider). Session persistence supports Redis, DynamoDB (AWS), Cosmos DB (Azure), and Firestore (GCP). Conversation threads support in-memory, Redis, Valkey, DynamoDB (AWS), Firestore (GCP), and Cosmos DB (Azure) backends. Generates configuration and code changes needed. license: Apache-2.0 metadata: author: yaalalabs version: "0.9.5" category: user --- # Add Capabilities Use this skill to enhance your Agent Kernel project with additional capabilities. ## Instructions for the Agent When the user wants to add a capability, follow this workflow: ### Step 1: Identify the Project Check for an existing Agent Kernel project with `pyproject.toml` and agent definition file. ### Step 2: Ask Which Capability Which capability would you like to add? 1. **Guardrails** — Content safety filters for input and/or output 2. **Tracing** — Observability and monitoring (Langfuse, OpenLLMetry, Pydantic Logfire, or AWS CloudWatch) 3. **Session Persistence** — Durable conversation state (Redis, DynamoDB, Cosmos DB, Firestore) 4. **Knowledge Base** — Durable cross-session knowledge tools (ChromaDB, Neo4j, Starburst, Open Knowledge Format markdown bundle, or custom backend) 5. **MCP Server** — Expose agents as Model Context Protocol tools 6. **A2A Server** — Agent-to-Agent communication protocol 7. **Hooks** — Custom pre/post processing (RAG, logging, prompt modification) 8. **Multimodal** — Image and file attachment support 9. **Conversation Threads** — Persistent, named conversation history keyed by `session_id` 10. **Sandbox** — Isolated code/command execution with pluggable providers, workload profiles, policy, and per-user identity 11. **AG-UI Server** — Stream any agent to an AG-UI-compliant frontend (text, tool calls, reasoning, shared state) 12. **Scheduled Tasks** — Deferred and recurring chat execution (a `schedule` block on a chat request, management routes, agent tools) 13. **Secret Resolution** — Read API keys through `SecretManager` (environment first, then AWS SSM Parameter Store or a custom provider) 14. **Human in the Loop** — Pause a run for a person's approval or answer and resume from their decision (a gated tool or an interrupt, a `resume` block on the next request) ### Step 3: Generate Changes --- #### Guardrails **Ask:** Input guardrails, output guardrails, or both? Which provider — OpenAI, AWS Bedrock, or Walled AI? **For OpenAI Guardrails:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api]>=0.9.5", # OpenAI guardrails use the openai extra — already included if using OpenAI framework ] ``` 2. Create `guardrails_input.json`: ```json [ { "type": "moderation", "moderation_config": { "content_type": ["violence", "sexual", "harassment", "self-harm"], "threshold": 0.5 } }, { "type": "jailbreak", "jailbreak_config": {} } ] ``` 3. Create `guardrails_output.json`: ```json [ { "type": "pii_detection", "pii_config": { "output_handling": "block", "entities": ["email_address", "phone_number", "ssn"] } } ] ``` 4. Update `config.yaml`: ```yaml guardrail: input: enabled: true type: openai model: gpt-4o-mini config_path: guardrails_input.json output: enabled: true type: openai config_path: guardrails_output.json ``` **For AWS Bedrock Guardrails:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,aws]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml guardrail: input: enabled: true type: bedrock guardrail_id: "" guardrail_version: "DRAFT" output: enabled: true type: bedrock guardrail_id: "" guardrail_version: "DRAFT" ``` 3. Prerequisites: Create a guardrail in AWS Bedrock Console and note the guardrail ID. **For Walled AI Guardrails:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,walledai]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml guardrail: input: enabled: true type: walledai pii: true # enable PII redaction on input output: enabled: true type: walledai pii: true # enable PII unmasking on output ``` 3. Set the environment variable: ```bash export WALLED_API_KEY="your-walledai-api-key" ``` 4. **How it works:** - **Input**: Text is checked for safety via `WalledProtect`. If unsafe, the request is blocked. If `pii: true`, text is redacted via `WalledRedact` and the PII mapping is stored in the session's non-volatile cache. - **Output**: If `pii: true`, redacted placeholders in the agent's reply are replaced with the original values using the stored mapping. --- #### Tracing (Observability) **Ask:** Which tracing backend — Langfuse, OpenLLMetry (Traceloop), Pydantic Logfire, or AWS CloudWatch? **For Langfuse:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,langfuse]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml trace: enabled: true type: langfuse ``` 3. Set environment variables: ```bash export LANGFUSE_PUBLIC_KEY="pk-..." export LANGFUSE_SECRET_KEY="sk-..." export LANGFUSE_HOST="https://cloud.langfuse.com" # or self-hosted URL ``` 4. No code changes needed — tracing is automatically applied to all agent executions. **For OpenLLMetry (Traceloop):** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,openllmetry]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml trace: enabled: true type: openllmetry ``` 3. Set environment variables per the Traceloop documentation. **For Pydantic Logfire:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,logfire]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml trace: enabled: true type: logfire ``` 3. Set the write token (optional — without it, Logfire runs locally and does not ship traces): ```bash export LOGFIRE_TOKEN="your-write-token" ``` 4. No code changes needed — tracing is automatically applied to all agent executions. **For AWS CloudWatch:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,cloudwatch]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml trace: enabled: true type: cloudwatch ``` 3. Set the region; credentials come from the standard AWS chain (env, profile, SSO, or instance/task role): ```bash export AWS_REGION="us-east-1" # Optional: export through the CloudWatch agent or an OpenTelemetry collector instead of straight to X-Ray # export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT="http://localhost:4318/v1/traces" ``` 4. One-time AWS setup: enable CloudWatch Transaction Search in the account, and attach the `AWSXrayWriteOnlyAccess` managed policy to the role that runs the agent. 5. No code changes needed — tracing is automatically applied to all agent executions. --- #### Session Persistence **Ask:** Which backend — Redis, DynamoDB (AWS), Cosmos DB (Azure), or Firestore (GCP)? **For Redis:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,redis]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml session: type: redis cache: 256 # LRU cache size (optional, improves performance) redis: prefix: "ak::" # Key prefix for namespacing url: "redis://localhost:6379" ttl: 3600 # Session TTL in seconds (optional) ``` **For DynamoDB:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,aws]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml session: type: dynamodb cache: 256 dynamodb: table_name: "" region: "us-east-1" ttl: 3600 ``` 3. Create a DynamoDB table with partition key `session_id` (String) and sort key `key` (String). Enable TTL on `expiry_time` attribute. **For Cosmos DB:** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,azure]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml session: type: cosmosdb cache: 256 cosmosdb: endpoint: "https://.table.cosmos.azure.com:443/" table_name: "" ttl: 3600 ``` 3. Set `AZURE_COSMOS_KEY` environment variable. **For Firestore (GCP):** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,gcp]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml session: type: firestore cache: 256 firestore: collection_name: "ak_sessions" project_id: "" # optional, inferred from ADC if omitted ttl: 604800 ``` 3. Enable a TTL policy on the Firestore collection pointing to the `expiry_time` field for automatic document expiry. --- #### Knowledge Base Add durable knowledge tools that your agents can query and update across sessions. **Ask:** Which backend do you want to use? - ChromaDB (semantic/vector) - Neo4j (graph/relationships) - Starburst (read-only SQL via Trino) - Open Knowledge Format bundle (a directory of markdown documents, on disk or in S3 — no database to run) - Custom adapter (developer extension) **1. Update `pyproject.toml` dependencies based on backend:** ```toml dependencies = [ "agentkernel[openai,api,chromadb]>=0.9.5", # for Chroma # or "agentkernel[openai,api,neo4j]>=0.9.5" # or "agentkernel[openai,api,trino]>=0.9.5" # OKF needs NO extra - pyyaml is a core dependency: # "agentkernel[openai,api]>=0.9.5" # ...unless the bundle is served from S3, which uses the aws extra: # "agentkernel[openai,api,aws]>=0.9.5" ] ``` **2. Configure backend + `KnowledgeBuilder` in your agent file (OpenAI example):** ```python from agents import Agent from agentkernel.cli import CLI from agentkernel.knowledgebase.chroma import ChromaManager from agentkernel.knowledgebase.knowledgebuilder import KnowledgeBuilder from agentkernel.openai import OpenAIModule, OpenAIToolBuilder backend = ChromaManager(name="ChromaDB").add_schema( { "description": "Semantic vector store for unstructured facts", "store_payload": {"text": "string", "source": "string"}, "read_payload": {"query": "string", "limit": "int"}, } ) kb = KnowledgeBuilder([backend]) kb_tools = kb.build() # -> get_schemas, read_kb, write_kb, get_all_kb_descriptions # (+ search_kb / fetch_kb / browse_kb when a registered # backend declares those capabilities - see step 4) # build(writable=False) omits write_kb for a read-only agent router = Agent( name="kb_router", instructions="Call get_schemas() first, then route reads/writes to the correct backend.", tools=OpenAIToolBuilder.bind(kb_tools), ) OpenAIModule([router]) if __name__ == "__main__": CLI.main() ``` **3. Multi-backend routing with semantic placeholders (for Starburst or mixed backends):** ```python kb = KnowledgeBuilder( [backend_a, backend_b], semantic_map={ "": "TABLE(kb_sheets.system.sheet(id => 'SHEET_ID'))", "": "mongodb.default.clients", }, ) ``` **4. Backend notes:** - `ChromaManager`: semantic search and fuzzy retrieval. Declares `search` + `writable`. - `Neo4jManager`: entity/relationship graphs and Cypher queries. Declares `query` + `writable`. - `StarburstManager`: **read-only** - it declares `writable=False`, so `write_kb` reports it as read-only instead of writing. Use `read_kb`; there is no way to mis-route a write into silent data loss. - `OKFManager`: an Open Knowledge Format bundle. Declares `search` (lexical), `fetch`, `browse` and `derives_schema`, and inherits the store's writability. **Which tools the agent gets is decided by those declarations.** Four are always built (`get_schemas`, `read_kb`, `write_kb`, `get_all_kb_descriptions`), unless `build(writable=False)` drops `write_kb` for an agent that must never write. `fetch_kb`, `browse_kb` and `search_kb` are added when a registered backend declares `fetch` / `browse` / `search`. So a Neo4j or Starburst app gets four tools, a Chroma app five, and an OKF app seven — the agent's prompt never names an operation nothing can serve. **4a. Open Knowledge Format bundle:** ```python from agentkernel.knowledgebase import DocumentStore, KnowledgeBuilder, LocalDocumentStore, OKFManager # A directory of markdown documents with YAML frontmatter. Read-only because a bundle checked # into git is a knowledge source, not a scratchpad; store writability folds into the backend's # declaration, so this one keyword is what makes write_kb report it as read-only. backend = OKFManager( LocalDocumentStore("./bundle", writable=False), name="OKF", description="Markdown knowledge bundle: one concept per topic, in browsable namespaces.", ) # Same bundle from S3 - a store swap, not a backend change (needs the aws extra): # backend = OKFManager(DocumentStore.from_uri("s3://my-bucket/bundles/kb"), name="OKF") kb_tools = KnowledgeBuilder([backend]).build() ``` There is **no `add_schema()` call**: `OKFManager` declares `derives_schema=True` and answers `get_schemas()` from the bundle itself. Tell the agent to `browse_kb` a namespace, then `fetch_kb` the concept path it saw — only `fetch_kb` returns a full document body. See `examples/cli/knowledgebase/openai/okf/` for a runnable demo with a checked-in bundle. **5. Environment variables (examples):** ```bash # Neo4j export NEO4J_URI="bolt://localhost:7687" export NEO4J_USERNAME="neo4j" export NEO4J_PASSWORD="password" # Starburst export STARBURST_HOST=".trino.galaxy.starburst.io" export STARBURST_USER="" export STARBURST_PASSWORD="" export STARBURST_PORT=443 ``` An OKF bundle needs no credentials when it is a local directory. Serving one from S3 uses the standard AWS chain (`AWS_REGION` plus whatever credentials the environment already provides), and the bundle location is a single string you can pass through `DocumentStore.from_uri()` — a bare path, `file://`, or `s3://bucket/prefix` — so one configuration value covers local-in-development and S3-in-production. **6. If the user asks for a new backend adapter:** - Add a custom backend by implementing `KnowledgeBase` under `ak-py/src/agentkernel/knowledgebase/`. - Use developer skill `.agents/skills/ak-dev-new-knowledgebase-integration/SKILL.md` for contributor workflows. --- #### MCP Server Expose your agents as MCP (Model Context Protocol) tools so other AI systems can discover and call them. 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,mcp]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml mcp: enabled: true expose_agents: true # Auto-expose all registered agents # agents: # Or list specific agents # - name: general # - name: math url: "http://localhost:8000" # Base URL of your server ``` 3. No code changes needed. The MCP server endpoint is automatically available. --- #### A2A Server Enable Agent-to-Agent communication via Google's A2A protocol. 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,a2a]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml a2a: enabled: true agents: - general # List agents to expose via A2A url: "http://localhost:8000" ``` 3. No code changes needed. The A2A well-known endpoint is automatically available at `/.well-known/agent.json`. --- #### AG-UI Server Stream any streaming-capable agent (OpenAI Agents SDK, LangGraph, Google ADK, Pydantic AI — not CrewAI or Smolagents) to a frontend over the [AG-UI protocol](https://github.com/ag-ui-protocol/ag-ui): text, tool calls, reasoning, and an optional shared JSON state, all as one typed event stream. 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,agui]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml agui: agents: ["general"] # omitted = every streaming-capable agent is reachable prefix: "/agui" default_agent: "general" # also serves POST /agui; must be one of `agents` when both are set state: enabled: true # attaches get_agui_state / update_agui_state client_context: enabled: true # attaches read-only get_forwarded_props / get_agui_context ``` 3. Mount `AGUIRequestHandler` with an `Authoriser` (or `AuthValidator`) — AG-UI has no anonymous mode, because a run executes an agent on the caller's behalf: ```python from agentkernel.agui import AGUIRequestHandler from agentkernel.api import RESTAPI from agentkernel.auth import Authoriser class MyAuthoriser(Authoriser): def authorise(self, token: str) -> str | None: ... # validate the token, return the caller's user_id or None RESTAPI.run(handlers=[AGUIRequestHandler(authoriser=MyAuthoriser())]) ``` 4. Routes are served under `agui.prefix`: `GET {prefix}/agents`, `POST {prefix}/{agent_name}`, and `POST {prefix}` when `default_agent` is set. See `examples/api/agui` for a full demo including a React/Vite frontend. --- #### Custom Hooks Add custom pre/post processing to your agents. **Pre-hook example (RAG injection):** ```python from agentkernel.core import Agent, PreHook, Session from agentkernel.core.model import AgentReply, AgentRequest, AgentRequestText class RAGPreHook(PreHook): async def on_run( self, session: Session, agent: Agent, requests: list[AgentRequest] ) -> list[AgentRequest] | AgentReply: # Extract the user's prompt prompt = "" for req in requests: if isinstance(req, AgentRequestText): prompt = req.prompt break # Retrieve relevant context (your RAG logic here) context = await self._retrieve_context(prompt) # Modify the prompt with additional context if context: enhanced_prompt = f"Context: {context}\n\nUser question: {prompt}" return [AgentRequestText(prompt=enhanced_prompt)] return requests async def _retrieve_context(self, query: str) -> str: # Implement your retrieval logic return "" def name(self) -> str: return "rag_prehook" ``` **Post-hook example (disclaimer):** ```python from agentkernel.core import Agent, PostHook, Session from agentkernel.core.model import AgentReply, AgentRequest class DisclaimerPostHook(PostHook): async def on_run( self, session: Session, requests: list[AgentRequest], agent: Agent, agent_reply: AgentReply ) -> AgentReply: agent_reply.response += "\n\n_Disclaimer: This is AI-generated content._" return agent_reply def name(self) -> str: return "disclaimer_posthook" ``` **Attach hooks to agents:** ```python from agentkernel.openai import OpenAIModule from agents import Agent agent = Agent(name="general", instructions="...") module = OpenAIModule([agent]) module.pre_hook(agent, [RAGPreHook()]) module.post_hook(agent, [DisclaimerPostHook()]) ``` **Framework-native run options:** Agent Kernel hooks wrap the whole run. For the framework's own per-run arguments and lifecycle hooks, declare them per agent with `module.run_options(agent, **options)`, chained like `pre_hook` / `post_hook`; repeated calls merge, the later call winning per key. Every keyword is one of the framework's native run arguments and Agent Kernel merges it into the call with the keys it owns written last. A native hook's callbacks run inside the Agent Kernel run, so `Session.current()` resolves in them (keep per-run state in the session's volatile cache, not on the hook instance, which is shared by concurrent runs). This works in any execution mode, including `rest_sync`; the streaming hook below is the framework-agnostic path for `stream` mode only. ```python from agents import RunConfig, RunHooks class ProgressHooks(RunHooks): async def on_tool_start(self, context, agent, tool) -> None: cache = Session.current().get_volatile_cache() cache.set("tool_calls", (cache.get("tool_calls") or 0) + 1) module.run_options( agent, max_turns=25, # the SDK default is 10 hooks=ProgressHooks(), run_config=RunConfig(call_model_input_filter=trim_history), ) ``` **Options computed per run:** pass a callable before the keywords, `module.run_options(agent, options_for, **static)`. It is called as `options_for(agent, session, requests)` on every run (sync or async), after the pre-hooks and inside the run, so `Session.current()` and `ToolContext.get()` resolve in it (except on Google ADK, where the tool context does not exist yet and `ToolContext.get()` raises), and its mapping is merged over the static keywords, a factory key winning. A reserved key in the result is rejected on the run, and a factory that raises fails that run like a framework error. One factory per agent, shared by concurrent runs: keep per-run state in the session. ```python def options_for(agent, session, requests): options = {"run_config": RunConfig(trace_metadata={"session_id": session.id})} if session.id.startswith("guest"): options["max_turns"] = 10 # over the static 25 below return options module.run_options(agent, options_for, max_turns=25, hooks=ProgressHooks()) ``` | Framework | `run_options` keywords go to | Turn limit | Progress hook | Reserved (raise at declaration) | |-----------|------------------------------|------------|---------------|---------------------------------| | OpenAI Agents SDK | `Runner.run` / `run_streamed` | `max_turns` | `hooks=RunHooks()` | `starting_agent`, `input`, `session`, `context`, `conversation_id`, `previous_response_id`, `auto_previous_response_id` | | LangGraph | `ainvoke` / `astream_events` (`config` is deep-merged: `configurable.thread_id` stays the session id, `callbacks` lists concatenate) | `config["recursion_limit"]` | `config["callbacks"]` | `input`, `version`, `stream_mode`, `output_keys`, `print_mode`, `config.configurable.thread_id`, and any `RunnableConfig` key at the top level | | Google ADK | the per-run `App` (`plugins`), the per-run `Runner(...)` constructor (services) and `run_async` (`run_config`; copied with `streaming_mode=SSE` in stream mode) | `RunConfig(max_llm_calls=...)` | `plugins=[BasePlugin()]` | `agent`, `app`, `app_name`, `node`, `session_service`, `auto_create_session`, `user_id`, `session_id`, `new_message`, `state_delta`, `invocation_id`, `yield_user_message` | | Pydantic AI | `agent.run` / `run_stream_events` (`event_stream_handler` is dropped in stream mode with one warning) | `UsageLimits(request_limit=...)` | `event_stream_handler` | `user_prompt`, `message_history`, `deps` | | CrewAI | the per-run `Crew(...)` constructor (`verbose=False` is an overridable default; agents resolve by `role`; `max_rpm` is a forwarded rate limit) | `max_iter` on the native `Agent` (needs nothing from Agent Kernel) | `step_callback` / `task_callback` | `agents`, `tasks`, `memory` | | smolagents | `agent.run` | `max_steps` | `step_callbacks` on the agent constructor (needs nothing from Agent Kernel) | `task`, `reset`, `additional_args`, `stream`, `return_full_result` | Worked demos: `examples/cli/-run-options` for each of the six frameworks, and `examples/cli/openai-dynamic-run-options` for options computed per run. **Streaming event hook (optional):** override `on_stream_event` on a `PostHook` to inspect or modify every event a streamed run produces while `execution.mode: stream` is active. Unlike `on_run` it sees the whole stream — message and reasoning text, tool call names, arguments and results, and the boundaries that pair them. Return the event to pass it on, a modified event of the same `type` to rewrite it, `None` to drop it, or a list to emit several events in its place (a list is emitted as-is and ends the chain for that event, so `return event` and `return [event]` differ). Raise `StreamHalt` to end the run: Agent Kernel closes any open boundary, emits one error chunk, and does not store the session. Only called when streaming; regular `on_run()` still handles the non-streaming path. ```python class RedactingPostHook(DisclaimerPostHook): async def on_stream_event(self, session, requests, agent, event): if event.type == "tool_call_result": return event.model_copy(update={"content": event.content.replace("SECRET", "***")}) return event ``` To rewrite text that spans fragments, hold each `text_delta` by returning `None` while accumulating it in `session.get_volatile_cache()`, then return `[TextDelta(...), event]` at `message_end`. Accumulate in the volatile cache, never on `self` — one hook instance serves every concurrent request. **Per-run framework context (optional):** hooks are the supported surface for the reserved `framework_context` session key — a framework-agnostic, picklable context/state dict that the runner injects into the native framework call (`context=`, `deps=`, session state, ...) and writes back after a successful run. It is never auto-created; seed it explicitly from a pre-hook, and read it back from a post-hook once the run has written its results: ```python class SeedCart(PreHook): async def on_run(self, session, agent, requests): if session.get_framework_context() is None: session.set_framework_context({"cart": []}) return requests def name(self): return "SeedCart" class AppendCart(PostHook): async def on_run(self, session, requests, agent, agent_reply): cart = (session.get_framework_context() or {}).get("cart", []) agent_reply.response += f"\n\nCurrent cart: {', '.join(cart) or '(empty)'}" return agent_reply def name(self): return "AppendCart" ``` Use `session.get_framework_context()` / `set_framework_context(dict)` / `clear_framework_context()` — these accessors are for hooks only. Tools must use their framework's native handle instead (`RunContextWrapper.context` on OpenAI, `RunContext.deps` on Pydantic AI, `tool_context.state` on ADK, ...) since a tool writing through `ToolContext.get().session` writes to a different object than the one the run is carrying. Round-trip fidelity is framework-dependent (full for OpenAI/Pydantic AI, partial for ADK/smolagents/LangGraph, unsupported for CrewAI) — see the framework's page under `docs/docs/frameworks/` for specifics. **Accessing the framework-native session (optional):** unlike `framework_context` above (an app-defined dict you seed yourself), `session.get_framework_session()` reaches the framework adapter's **own** session object directly — the same live object each runner stores under its runner-name key (e.g. `"openai"`), without you needing to name that key. It only works from inside a hook or a tool (an agent must currently be running — it raises `RuntimeError` otherwise), and mutating the returned object through its own methods is visible immediately, no `session.set(...)` needed: ```python class HistoryTrimHook(PostHook): async def on_run(self, session, requests, agent, agent_reply): openai_session = session.get_framework_session() if openai_session is not None: items = await openai_session.get_items() if len(items) > 20: await openai_session.clear_session() await openai_session.add_items(items[-20:]) # keep only the most recent 20 return agent_reply def name(self): return "HistoryTrimHook" ``` See `examples/cli/session-context/hooks.py` (`HistoryTrimHook`) for a complete example that caps the OpenAI Agents SDK's raw conversation history after every turn. --- #### Multimodal Support Enable image and file processing in your agents. **Ask:** Which storage backend — in-memory (default, dev), Redis (production), or DynamoDB (serverless/AWS)? **Basic setup (in-memory storage, good for development):** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,multimodal]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml multimodal: enabled: true max_attachments: 20 description_model: "gpt-4o" # Model for generating brief descriptions analysis_model: "gpt-4o" # Model for detailed analysis via tool ``` 3. No further code changes needed. When enabled: - A system tool (`analyze_attachments`) is automatically attached to all agents - Image/file attachments in requests are processed, described by a vision LLM, and stored in a separate in-memory attachment store (outside the conversation history/session state, non-persistent, in-process) - Binary data is kept out of conversation history to prevent memory bloat - Agents see attachment IDs and descriptions in their context - Agents can call `analyze_attachments(attachment_ids, prompt)` for detailed analysis **For Redis storage (production, persistent, distributed):** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,redis,multimodal]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml multimodal: enabled: true storage_type: redis max_attachments: 20 description_model: "gpt-4o" analysis_model: "gpt-4o" redis: url: "redis://localhost:6379" prefix: "ak:attachments:" ttl: 604800 # Attachment TTL in seconds (7 days) ``` **For DynamoDB storage (serverless/AWS):** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,aws,multimodal]>=0.9.5", ] ``` 2. Update `config.yaml`: ```yaml multimodal: enabled: true storage_type: dynamodb max_attachments: 20 description_model: "gpt-4o" analysis_model: "gpt-4o" dynamodb: table_name: "ak-attachments" ttl: 604800 ``` 3. Create a DynamoDB table with partition key `session_id` (String) and sort key `attachment_id` (String). Enable TTL on `expiry_time` attribute. **Storage type comparison:** | Type | Persistence | Multi-process | Setup | Best for | |------|-------------|---------------|-------|----------| | `in_memory` | Lost on restart | Single process | None | Dev/testing | | `redis` | Persistent | Distributed | Redis server | Production | | `dynamodb` | Persistent | Distributed | AWS table | Serverless/Lambda | **How it works:** - When a user sends an image or file, a vision LLM generates a brief one-sentence description - The binary data is saved to the configured storage backend (not in conversation history) - The agent receives the text prompt enriched with attachment IDs and descriptions - When the agent needs to inspect an attachment in detail, it calls the `analyze_attachments` tool - The tool retrieves the binary from storage, sends it to the analysis LLM, and returns clean text **Send multimodal requests via API:** ```bash # Base64 image curl -X POST http://localhost:8000/run \ -H "Content-Type: application/json" \ -d '{ "prompt": "What is in this image?", "session_id": "test-1", "agent": "general", "image": "", "image_name": "photo.jpg" }' # File upload (multipart) curl -X POST http://localhost:8000/run \ -F "prompt=Summarize this document" \ -F "session_id=test-1" \ -F "agent=general" \ -F "file=@document.pdf" ``` **Environment variables (all storage types):** ```bash # Required: LLM API key for vision models export OPENAI_API_KEY="sk-..." # For Redis storage export AK_MULTIMODAL__STORAGE_TYPE=redis export AK_MULTIMODAL__REDIS__URL="redis://localhost:6379" # For DynamoDB storage export AK_MULTIMODAL__STORAGE_TYPE=dynamodb export AK_MULTIMODAL__DYNAMODB__TABLE_NAME="ak-attachments" ``` --- #### Conversation Thread Support Enable persistent, named conversation threads keyed by `session_id`. **Ask:** Which thread store backend — in-memory (default, dev), Redis, Valkey, DynamoDB (AWS), Firestore (GCP), or Cosmos DB (Azure)? **Basic setup (in-memory store, good for development):** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api]>=0.9.5", ] ``` 2. Mount the thread handler in the app; this is what enables the feature (it serves the standard chat routes with thread recording, plus the thread read routes): ```python from agentkernel.api import RESTAPI from agentkernel.thread import AgentThreadRequestHandler RESTAPI.run(handlers=[AgentThreadRequestHandler()]) ``` 3. Update `config.yaml` (selects the store backend; constructing the handler without this block fails fast at startup): ```yaml thread: type: in_memory # other supported backends: redis | valkey | dynamodb | firestore | cosmosdb ``` 4. When enabled: - `user_id` becomes required on the thread handler's chat requests (other surfaces are unaffected) - A thread is auto-created on a session's first request - `GET /api/v1/threads` and `GET /api/v1/threads/{session_id}` are served by the same handler for reading thread history (open by default, or protected by a pluggable `Authoriser`) - Threads are auto-named by an LLM call deriving a concise title from the first prompt (falls back to a truncated prompt prefix without `litellm`/an API key) - Sending `thread_name` on any chat request sets/renames the thread and locks it against automatic naming **For LLM-based thread naming**, add the `thread` extra: ```toml dependencies = [ "agentkernel[openai,api,thread]>=0.9.5", ] ``` ```yaml thread: type: in_memory naming: model: "gpt-4o-mini" # LiteLLM model used to name threads max_length: 80 ``` **For Redis storage (production, persistent, distributed):** ```toml dependencies = [ "agentkernel[openai,api,redis,thread]>=0.9.5", ] ``` ```yaml thread: type: redis redis: url: "redis://localhost:6379" prefix: "ak:thread:" ttl: 2592000 # Thread TTL in seconds (30 days, 0 disables) ``` **For Valkey storage (production, persistent, distributed, Redis-protocol compatible):** ```toml dependencies = [ "agentkernel[openai,api,valkey,thread]>=0.9.5", ] ``` ```yaml thread: type: valkey valkey: url: "valkey://localhost:6379" prefix: "ak:thread:" ttl: 2592000 # Thread TTL in seconds (30 days, 0 disables) ``` **For DynamoDB storage (serverless/AWS):** ```toml dependencies = [ "agentkernel[openai,api,aws,thread]>=0.9.5", ] ``` ```yaml thread: type: dynamodb dynamodb: table_name: "ak-agent-threads" # partition key session_id (S), sort key sk (S) ttl: 0 ``` **For Firestore storage (serverless/GCP):** ```yaml thread: type: firestore firestore: collection_name: "ak-agent-threads" ttl: 0 ``` **For Cosmos DB storage (Azure, Table API):** ```yaml thread: type: cosmosdb cosmosdb: connection_string: "${AZURE_COSMOS_CONNECTION_STRING}" table_name: "akagentthreads" ``` **Protecting the read endpoints with an Authoriser:** ```python from typing import Optional from agentkernel.api import RESTAPI from agentkernel.auth import Authoriser from agentkernel.thread import AgentThreadRequestHandler class DemoAuthoriser(Authoriser): def authorise(self, token: str) -> Optional[str]: # Validate the ****** against your own auth provider, return the user_id or None. return {"alice-token": "alice", "bob-token": "bob"}.get(token) RESTAPI.run(handlers=[AgentThreadRequestHandler(authoriser=DemoAuthoriser())]) ``` With an Authoriser configured, thread listings are scoped to the resolved `user_id` and reading another user's thread is rejected (403). Without one, the read routes are open. **Attachments in thread mode:** require `multimodal.enabled: true` with a shared attachment store — `in_memory`, `redis`, or `dynamodb` (`session_cache` is rejected). **Send a chat request with a thread:** ```bash curl -X POST http://localhost:8000/api/v1/chat \ -H "Content-Type: application/json" \ -d '{"prompt": "What is the capital of France?", "session_id": "ses-1", "user_id": "alice", "thread_name": "Capitals quiz"}' ``` See `examples/api/thread-openai` and `examples/api/multimodal/thread-openai`. --- #### Scheduled Tasks **What it does:** Lets a chat request run later, once or repeatedly. A request carrying a `schedule` block is not executed — it is registered as a scheduled task and acknowledged with **HTTP 202**. When an occurrence is due, the provider delivers the stored prompt into the input queue as a plain chat request and the normal execution path runs it. The block also injects five agent tools (`create_schedule`, `list_schedules`, `get_schedule`, `update_schedule`, `delete_schedule`) so the agent can defer work itself. The management routes are **not** mounted from config — the application mounts `ScheduleRESTRequestHandler` when it wants them, exactly as it mounts the Slack and thread handlers. **Ask:** Which provider — `local` (in-process thread, development) or `eventbridge` (AWS EventBridge Scheduler, production)? And which task store — in-memory (default, dev), Redis, Valkey, or DynamoDB (AWS)? **Important:** occurrences are delivered *into the input queue*, so scheduling requires the queue execution pipeline. Locally the `in_memory` transport satisfies this inside one process; on AWS it means deploying in queue mode. **Basic setup (local provider, in-memory store — development):** 1. Update `pyproject.toml`: ```toml dependencies = [ "agentkernel[openai,api,cron]>=0.9.5", ] ``` The `cron` extra brings `croniter`, needed for cron parsing. 2. Update `config.yaml`. The presence of the `schedule` block is what enables deferring and the agent tools; the management routes are mounted by the app in step 3: ```yaml schedule: provider: type: local # other supported providers: eventbridge store: type: in_memory # other supported backends: redis | valkey | dynamodb # agents: [assistant] # restrict the schedule tools to named agents; omitted = all agents execution: mode: rest_sync queues: type: in_memory # required: the provider fires occurrences into the input queue ``` 3. Mount the management routes in `app.py`. Nothing is mounted from config, so an app that skips this step still defers requests and still gets the agent tools — it just serves no `/api/v1/schedules` routes: ```python from agentkernel.pipeline import IOHandler from agentkernel.schedule import ScheduleRESTRequestHandler if __name__ == "__main__": # config.yaml selects the in_memory queue transport, so this boots the whole single-process # pipeline. The passed handlers are mounted alongside the pipeline's own chat route. IOHandler.run(handlers=[ScheduleRESTRequestHandler()]) ``` 4. When enabled: - A JSON chat request may carry a `schedule` block: exactly one of `at` (ISO-8601 local wall-clock timestamp, must be in the future) or `cron` (standard 5-field expression), plus `timezone` (IANA, default `UTC`) and `session_mode` (`reuse` the originating session, or `new` for a fresh session per occurrence) - `user_id` becomes **required** on any request that schedules: it is the owner the task is stored under and the identity later reads and changes are checked against - `GET /api/v1/schedules` (cursor-paginated) and `GET`/`PUT`/`DELETE /api/v1/schedules/{task_id}` are mounted for listing, reading, amending and cancelling (open by default, or protected by a pluggable `Authoriser`). There is deliberately no `POST` — creation is the chat block or the agent tool - `PUT` is full-replacement: send every value, including the ones that are not changing. `status` covers the `active`/`paused` switch; a cancelled task keeps its record as the audit trail - The five schedule tools and their guidance are injected into every agent's system prompt; each acts as the invoking user, so an agent can never reach another user's schedules - A scheduled request creates no conversation thread; the occurrences that later fire do - Multipart chat routes cannot carry a `schedule` block — use the JSON route **Send a chat request with a schedule:** ```bash # One-time curl -i -X POST http://localhost:8000/api/v1/chat \ -H "Content-Type: application/json" \ -d '{"prompt": "Send me the daily summary", "session_id": "ses-1", "user_id": "alice", "schedule": {"at": "2030-01-31T09:00:00", "timezone": "Asia/Colombo"}}' # HTTP/1.1 202 Accepted # {"result":"{\"status\": \"SCHEDULED\", \"scheduled_task_id\": \"74ca19a5-...\", \"session_id\": \"ses-1\"}","session_id":"ses-1"} # Recurring, each occurrence in a fresh session curl -X POST http://localhost:8000/api/v1/chat \ -H "Content-Type: application/json" \ -d '{"prompt": "Send the weekly report", "session_id": "ses-2", "user_id": "alice", "schedule": {"cron": "0 9 * * 1", "timezone": "Asia/Colombo", "session_mode": "new"}}' # Manage curl "http://localhost:8000/api/v1/schedules?user_id=alice" curl -X DELETE http://localhost:8000/api/v1/schedules/{task_id} ``` **For production on AWS (EventBridge Scheduler + DynamoDB):** ```toml dependencies = [ "agentkernel[openai,api,aws,cron]>=0.9.5", ] ``` ```yaml schedule: provider: type: eventbridge store: type: dynamodb dynamodb: table_name: "ak-agent-schedules" # partition key task_id (S), no sort key ttl: 0 # 0 disables expiry (the default) ``` `group_name`, `role_arn` and `queue_arn` under `schedule.provider.eventbridge` are supplied by the Terraform modules as `AK_SCHEDULE__PROVIDER__EVENTBRIDGE__*` environment variables — do not hardcode them. See the `ak-cloud-deploy` skill. **For Redis or Valkey task storage:** ```toml dependencies = [ "agentkernel[openai,api,redis,cron]>=0.9.5", # or valkey ] ``` ```yaml schedule: store: type: redis # or valkey, with a `valkey:` block redis: url: "redis://localhost:6379" prefix: "ak:schedule:" ttl: 0 # unlike threads this defaults to 0 — an expired task would stop firing silently ``` **Topology rules (validated at startup, not at first use):** | Combination | Rejected because | |---|---| | `local` provider + a broker transport (`sqs`/`kafka`/`nats`) | The in-process timers are unreachable from the process serving the management routes — a cancellation would report success while the timer kept firing | | `local` provider + a shared store | Same split: the timers and the records must live together | | `in_memory` store + a broker transport | The records would be split across the runner and IO-handler processes | | `eventbridge` provider + a non-`sqs` transport | Delivery is baked into the schedule registration as an SQS target | **Protecting the management routes with an Authoriser:** the routes are mounted by the application, so pass the `Authoriser` to the `ScheduleRESTRequestHandler` constructor: ```python from typing import Optional from agentkernel.auth import Authoriser from agentkernel.pipeline import IOHandler from agentkernel.schedule import ScheduleRESTRequestHandler class DemoAuthoriser(Authoriser): def authorise(self, token: str) -> Optional[str]: # Validate the ****** against your own auth provider, return the user_id or None. return {"alice-token": "alice", "bob-token": "bob"}.get(token) if __name__ == "__main__": IOHandler.run(handlers=[ScheduleRESTRequestHandler(authoriser=DemoAuthoriser())]) ``` With an Authoriser configured, listings are scoped to the resolved `user_id` and reading or changing another user's schedule is rejected (403). Without one, the routes are open. See `examples/api/schedule-openai`. --- #### Human in the Loop **What it does:** Lets a run stop mid-way to ask a person something — approve this gated tool, answer this question — and resume from their decision minutes or hours later, possibly on another replica. There is **no `enabled` flag and no config block**: a pause happens when the framework decides one is needed, so it is declared where the framework declares it. Agent Kernel answers **HTTP 202** with `status: "PAUSED"` and the pending interruptions; the next request carries a `resume` block instead of a prompt. **Ask:** Which framework is the agent on? Only four can pause — OpenAI Agents SDK, LangGraph, Pydantic AI and Google ADK. CrewAI and smolagents report `supports_pause = False` rather than pretending. **Important:** the pause is written into the **session**, so a durable one needs a shared session backend. With `session.type: in_memory` the record lives in one process and the replica receiving the decision has never heard of the pause — Agent Kernel logs a warning the first time that happens. **Declaring it, per framework:** ```python # OpenAI Agents SDK — a gated tool @function_tool(needs_approval=True) def issue_refund(order_id: str) -> str: ... # LangGraph — an interrupt returns whatever the resume supplies def ask(state): choice = interrupt({"question": "Which method?", "options": ["Card", "Credit"]}) # Pydantic AI — needs DeferredToolRequests among its output types, or neither axis can pause agent = Agent(model="openai:gpt-4.1-mini", output_type=[str, DeferredToolRequests]) @agent.tool(name="ask_size") def ask_size(ctx: RunContext, q: str) -> str: raise CallDeferred # a value; `requires_approval=True` asks for a verdict instead # Google ADK — a result, or a verdict LongRunningFunctionTool(func=ask_address) FunctionTool(func=issue_refund, require_confirmation=True) ``` **Answering it:** ```bash # The pause: HTTP 202 # {"status": "PAUSED", "run_id": "9f2c...", "agent": "support", # "interruptions": [{"id": "call_abc123", "kind": "tool_call", "tool_name": "issue_refund"}]} curl -X POST http://localhost:8000/api/v1/chat -H "Content-Type: application/json" -d '{ "agent": "support", "session_id": "user-123", "resume": {"decisions": [{"id": "call_abc123", "status": "approved"}]} }' ``` A decision takes `status` (`approved` | `denied` | `cancelled` — "nobody decided", which never reaches the model as a refusal), an optional `message` for the human's own words, and an optional `payload` for a structured answer. `run_id` is optional: the run resolves from the interruption ids. **What each framework can carry back differs, and Agent Kernel refuses rather than silently dropping.** OpenAI takes no `payload` at all (an approval is a boolean); Pydantic AI takes one only as a JSON object on an approval, since it becomes the call's `override_args`; ADK refuses one on a confirmation and cannot tell `cancelled` from `denied` there. Point the user at the [Human in the Loop](https://kernel.yaala.ai/docs/advanced/human-in-the-loop) page for the full table before they design the question. #### Sandbox **What it does:** Lets agents execute code and shell commands in an isolated, permission-bounded environment. When enabled, agents automatically gain sandbox tools (`run_code`, `run_command`, `write_sandbox_file`, `read_sandbox_file`, `check_sandbox_task`, `list_sandbox_sessions`, `new_sandbox_session`, `destroy_sandbox_session`) and the usage guidance is injected into their system prompt — the agent's own instructions need not mention the sandbox. **Ask:** Which provider — `local_subprocess` (no isolation; dev/test only), `docker` (container isolation; needs the `sandbox-docker` extra and a Docker daemon), or another shipped provider (`kubernetes` pods, `e2b` micro-VMs, `daytona` cloud containers, `ec2_ssm` attach-only; see the [Sandbox guide](https://kernel.yaala.ai/docs/advanced/sandbox))? Should it apply to all agents or only some (the `agents` list)? For executions longer than the process can wait, the `queue` broker flavor runs them on a separate worker (`sandbox.broker.flavor: queue`; same guide). **1. Install the extra (docker only):** ```bash pip install "agentkernel[sandbox-docker]" ``` **2. Add a `sandbox` block to `config.yaml`.** Minimal (single-backend sugar synthesizes a `default` profile): ```yaml sandbox: enabled: true type: local_subprocess # or: docker local_subprocess: {} # or a docker: { image: python:3.12-slim } block broker: flavor: thread # thread (CLI/REST default) | embedded ``` With explicit profiles, policy, and scoping: ```yaml sandbox: enabled: true agents: [coder] # optional: only these agents get the sandbox (omit = all) default_profile: workspace tool_output_max_chars: 8000 profiles: workspace: type: docker scope: per_session # per_call | per_session | per_runtime environment: managed # managed (default) | attached (connect to an existing environment; needs attach_to) idle_timeout: 1800 policy: network_egress: deny # allow | deny | allowlist cpu: 1.0 memory_mb: 512 timeout: 30.0 strict: true # fail closed when the provider can't enforce a dimension docker: image: python:3.12-slim broker: flavor: thread ``` **3. No agent code changes needed** — the tools and prompt guidance attach automatically. Keep the agent's instructions about *what* to do; the sandbox usage is injected. :::caution `local_subprocess` runs code directly on the host with **no isolation** — dev/test only. Use `docker` (or another isolating provider) in production. ::: For per-user identity (running sandboxed code under the invoking user's identity), set `principal_resolver` to a dotted path and a profile's `identity.mode: user`; see the [Sandbox guide](https://kernel.yaala.ai/docs/advanced/sandbox) and the `examples/sandbox/identity` example. --- #### Secret Resolution **What it does:** `SecretManager.current().get("OPENAI_API_KEY")` resolves an environment-variable-style key (`^[A-Z][A-Z0-9_]*$`) in a fixed order: the process environment, then a TTL'd process cache, then the configured provider (`secret.provider.type`). A set, non-empty environment variable always wins; an empty one counts as absent. A resolved value is **never written back to `os.environ`** — hand it to the SDK explicitly. The capability is always on (no `enabled` flag); the default `env` provider costs nothing. **Ask:** Which provider — `env` (default; the process environment), `aws_ssm` (AWS SSM Parameter Store; needs the `aws` extra and `secret.prefix`), or a dotted path to your own `SecretProvider` subclass? **1. Install the extra (aws_ssm only):** ```bash pip install "agentkernel[aws]" ``` **2. Add a `secret` block to `config.yaml`** (optional for `env`, which is the default): ```yaml secret: provider: type: aws_ssm # env (default) | aws_ssm | my_pkg.secrets.MyProvider prefix: myproduct-dev-agents # required by aws_ssm; AWS Terraform injects AK_SECRET__PREFIX when ssm_enabled = true cache_ttl: 300 # seconds a provider hit is cached; 0 disables caching (rotation-pickup window) ``` With `aws_ssm`, the key `OPENAI_API_KEY` is read from the SSM parameter `/ak/{prefix}/openai_api_key` (key lowercased, `GetParameter` with decryption). `prefix` must be a single path segment. Create the parameters yourself as `SecureString`. **3. Resolve secrets in agent code instead of reading `os.environ`:** ```python from agentkernel.secret import SecretManager from agents import set_default_openai_key # Required secret: a miss raises SecretNotFoundError at startup, not on the first model call set_default_openai_key(SecretManager.current().get("OPENAI_API_KEY")) def get_weather(city: str) -> str: # Optional secret: default=None turns a miss into None instead of raising api_key = SecretManager.current().get("WEATHER_API_KEY", default=None) ... ``` Errors: `SecretNotFoundError` (no layer has the key and no `default`), `SecretError` (the provider failed — credentials, network, IAM; never masked by `default`), `ValueError` (malformed key). Use `SecretManager.current().invalidate(key)` / `.clear()` to drop cached values. A custom provider subclasses `agentkernel.secret.SecretProvider`, implements `get_secret(key) -> Optional[str]` (return `None` on a miss), and can be checked with the `agentkernel.secret.testing.SecretProviderContract` pytest suite. See `examples/cli/openai-secret` (env) and `examples/aws-serverless/openai` (aws_ssm). --- ### What to Do Next You've added new capabilities to your project. Here's what you might do next: - **Add more tools & agents** → Use the `ak-build` skill to add new tools and specialist agents that leverage your new capabilities (e.g., agents that use guardrails or hooks). - **Connect a messaging platform** → Use the `ak-add-integration` skill to add Slack, WhatsApp, Telegram, or other channels so users can interact with your enhanced agents. - **Deploy to cloud** → Use the `ak-cloud-deploy` skill to deploy your agent (with all its capabilities) to AWS or Azure. - **Set up testing** → Use the `ak-test` skill to verify your capabilities work correctly — especially guardrails and hooks. - **Extend knowledge backends** → Contributors can use `.agents/skills/ak-dev-new-knowledgebase-integration/SKILL.md` to add new KnowledgeBase adapters.