--- name: ai-observability description: > Use when adding Spring AI-specific model observations, token usage, latency, externally configured cost attribution, advisor telemetry, or protected prompt and completion logging. Use production-observability for general service metrics, health, logs, and OTLP setup. --- # AI Observability ## Dependencies ```xml org.springframework.boot spring-boot-starter-actuator io.micrometer micrometer-registry-prometheus ``` ## Spring AI Built-in Observability Spring AI 1.0+ includes built-in Micrometer instrumentation: ```yaml spring: ai: chat: observations: log-prompt: true # GA renamed include-prompt → log-prompt. OFF in prod (PII). log-completion: true # GA renamed include-completion → log-completion management: metrics: tags: application: order-service endpoints: web: exposure: include: health,prometheus,metrics ``` Auto-generated metrics (OpenTelemetry GenAI semantic conventions): - `gen_ai.client.operation` — model call latency, tagged with provider and model - `gen_ai.client.token.usage` — token counts (input/output/total) - `spring.ai.chat.client` — ChatClient-level operation timer/span ## Custom AI Metrics ```java @Component @RequiredArgsConstructor public class AiMetrics { private final MeterRegistry meterRegistry; private final Timer.Builder promptTimer = Timer.builder("ai.prompt.latency") .description("LLM prompt latency"); private final Counter.Builder tokenCounter = Counter.builder("ai.tokens.used") .description("Total tokens consumed"); public T track(String operation, String model, Supplier call) { return Timer.builder("ai.prompt.latency") .tag("operation", operation) .tag("model", model) .register(meterRegistry) .recordCallable(() -> call.get()); } public void recordTokens(String operation, String model, int inputTokens, int outputTokens) { Counter.builder("ai.tokens.used") .tag("operation", operation) .tag("model", model) .tag("type", "input") .register(meterRegistry) .increment(inputTokens); Counter.builder("ai.tokens.used") .tag("operation", operation) .tag("model", model) .tag("type", "output") .register(meterRegistry) .increment(outputTokens); } } ``` ## Prompt/Response Logging Advisor GA replaced the whole advisor API: `CallAroundAdvisor` → `CallAdvisor`, `AdvisedRequest` → `ChatClientRequest`, `AdvisedResponse` → `ChatClientResponse`, and `Usage.getGenerationTokens()` → `getCompletionTokens()`. Agents reliably generate the old one — it does not compile on 1.0. ```java @Component public class AiAuditAdvisor implements CallAdvisor { private static final Logger log = LoggerFactory.getLogger(AiAuditAdvisor.class); @Override public ChatClientResponse adviseCall(ChatClientRequest request, CallAdvisorChain chain) { String requestId = UUID.randomUUID().toString(); long start = System.currentTimeMillis(); log.info("[AI-AUDIT] requestId={} promptLength={}", requestId, request.prompt().getUserMessage().getText().length()); try { ChatClientResponse response = chain.nextCall(request); long latency = System.currentTimeMillis() - start; ChatResponse chatResponse = response.chatResponse(); if (chatResponse != null && chatResponse.getMetadata() != null) { Usage usage = chatResponse.getMetadata().getUsage(); log.info("[AI-AUDIT] requestId={} latencyMs={} inputTokens={} outputTokens={}", requestId, latency, usage.getPromptTokens(), usage.getCompletionTokens()); // GA: not getGenerationTokens() } return response; } catch (Exception e) { log.error("[AI-AUDIT] requestId={} FAILED after {}ms", requestId, System.currentTimeMillis() - start, e); throw e; } } @Override public String getName() { return "AiAuditAdvisor"; } @Override public int getOrder() { return Ordered.LOWEST_PRECEDENCE; } } ``` ## Cost attribution - Keep provider prices in externally managed configuration with an effective date and currency. - Key prices by the exact provider model identifier returned in usage metadata. - Reject an unknown model instead of silently applying a default price. - Preserve the raw token usage so historical costs can be recalculated after pricing changes. - Prefer provider billing exports for invoices; application estimates are operational signals only. ## Structured AI Audit Log (DB) ```java @Entity @Table(name = "ai_audit_log") public class AiAuditLog { @Id @GeneratedValue(strategy = GenerationType.UUID) private UUID id; private String operation; private String model; private int inputTokens; private int outputTokens; private double estimatedCostUsd; private long latencyMs; private boolean success; private Instant createdAt; } // Async to avoid blocking main flow @Async public void saveAuditLog(AiAuditLog log) { auditLogRepository.save(log); } ``` ## application.yml — Full Observability ```yaml management: endpoints: web: exposure: include: health,prometheus,metrics,info metrics: distribution: percentiles-histogram: ai.prompt.latency: true # enables P50/P95/P99 tracing: sampling: probability: 1.0 # 100% trace sampling in dev, reduce in prod logging: level: org.springframework.ai: DEBUG # enable in dev only ``` ## Gotchas - Agent implements `CallAroundAdvisor`/`AdvisedRequest` — removed in GA; use `CallAdvisor`/`ChatClientRequest` - Agent calls `usage.getGenerationTokens()` — GA renamed it to `getCompletionTokens()` - Agent logs full prompts in production — keep `log-prompt: false` for PII safety - Agent skips async on audit saves — always `@Async` to avoid latency impact, and put the `@Async` method on a **separate bean**; calling it on `this` bypasses the proxy and runs synchronously - Agent hardcodes token pricing — extract to config, prices change - Agent misses failed calls in metrics — track errors separately with error tag