{"id":"voiceagentbench","name":"VoiceAgentBench","primary_focus":"End-to-end speech-based agents on realistic tool-driven tasks","audio_input":"yes","audio_output":"no","multi_turn":"yes","tool_use":"yes","goal_completion":"partial","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"no","public_data":"yes","code_license":"Krutrim Community License Agreement 1.0","data_license":"other","github_url":"https://github.com/ola-krutrim/VoiceAgentBench","huggingface_url":"https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench","paper_url":"https://arxiv.org/abs/2510.07978","last_verified":"2026-09-01","evidence_notes":"Public code and audio-backed data cover single, parallel, sequential, multi-turn, and safety tool-call subsets. Scoring emphasizes tool arguments and refusals rather than live spoken output or computer control."} {"id":"audio2tool","name":"Audio2Tool","primary_focus":"Spoken tool-call selection and argument extraction across smart-home, wearable, and vehicle tasks","audio_input":"yes","audio_output":"no","multi_turn":"partial","tool_use":"yes","goal_completion":"no","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"no","public_data":"yes","code_license":"not stated","data_license":"CC-BY-NC-4.0","github_url":null,"huggingface_url":"https://huggingface.co/datasets/RVtech/Audio2Tool","paper_url":"https://arxiv.org/abs/2604.22821","last_verified":"2026-09-01","evidence_notes":"The public dataset contains 16,843 queries and 36,421 audio files across eight tiers, including multi-intent, correction, multi-turn, and overlapping-speech cases. It evaluates audio-to-tool-call generation, not execution against a stateful environment or spoken-response quality."} {"id":"eva","name":"EVA","primary_focus":"End-to-end conversational voice-agent accuracy and experience","audio_input":"yes","audio_output":"yes","multi_turn":"yes","tool_use":"yes","goal_completion":"yes","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"yes","public_data":"yes","code_license":"MIT","data_license":"MIT","github_url":"https://github.com/ServiceNow/eva","huggingface_url":"https://huggingface.co/datasets/ServiceNow-AI/eva-bench","paper_url":"https://arxiv.org/abs/2605.13841","last_verified":"2026-09-01","evidence_notes":"Bot-to-bot evaluations cover 213 scenarios across three enterprise domains, complete multi-turn spoken conversations, task accuracy, interaction experience, voice perturbations, and 12 evaluated systems."} {"id":"voiceassistant-eval","name":"VoiceAssistant-Eval","primary_focus":"Listening, speaking, viewing, multi-turn behavior, and safety for general voice assistants","audio_input":"yes","audio_output":"yes","multi_turn":"yes","tool_use":"no","goal_completion":"no","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"no","public_data":"yes","code_license":"not stated","data_license":"MIT","github_url":"https://github.com/mathllm/VoiceAssistant-Eval","huggingface_url":"https://huggingface.co/datasets/MathLLMs/VoiceAssistant-Eval","paper_url":"https://arxiv.org/abs/2509.22651","last_verified":"2026-09-01","evidence_notes":"A broad multimodal voice-assistant suite with public examples across listening, speaking, viewing, multi-turn interaction, and safety. It does not primarily evaluate tool execution or real-world task completion."} {"id":"voicebench","name":"VoiceBench","primary_focus":"Multi-faceted instruction following, knowledge, reasoning, and safety for LLM-based voice assistants","audio_input":"yes","audio_output":"no","multi_turn":"partial","tool_use":"no","goal_completion":"no","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"no","public_data":"yes","code_license":"Apache-2.0","data_license":"Apache-2.0","github_url":"https://github.com/MatthewCYM/VoiceBench","huggingface_url":"https://huggingface.co/datasets/hlt-lab/voicebench","paper_url":"https://arxiv.org/abs/2410.17196","last_verified":"2026-09-01","evidence_notes":"Public audio subsets cover open-ended and multiple-choice QA, instruction following, reasoning, safety, human-recorded speech, and a 46-example multi-turn subset. Evaluation consumes spoken prompts but scores assistant response content rather than spoken-output quality or tool execution."} {"id":"voicecomputerbench-talkact","name":"VoiceComputerBench / TalkAct","primary_focus":"Real-time voice conversation while operating browser-based computer tasks","audio_input":"yes","audio_output":"yes","multi_turn":"yes","tool_use":"yes","goal_completion":"yes","computer_or_browser_action":"yes","meeting_or_long_form":"no","real_time":"yes","public_data":"yes","code_license":"MIT","data_license":"MIT","github_url":"https://github.com/19PINE-AI/TalkAct","huggingface_url":null,"paper_url":"https://github.com/19PINE-AI/TalkAct/blob/main/paper/paper_arxiv.pdf","last_verified":"2026-09-01","evidence_notes":"VoiceComputerBench evaluates a simulated phone caller and an agent that talks while acting through Playwright on hermetic browser sites."} {"id":"tau2-bench","name":"tau2-bench / tau-Voice","primary_focus":"Tool-agent-user interaction in realistic domains","audio_input":"yes","audio_output":"yes","multi_turn":"yes","tool_use":"yes","goal_completion":"yes","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"yes","public_data":"yes","code_license":"MIT","data_license":"MIT","github_url":"https://github.com/sierra-research/tau2-bench","huggingface_url":null,"paper_url":"https://arxiv.org/abs/2603.13686","last_verified":"2026-09-01","evidence_notes":"The public framework covers tool use, dynamic user interaction, domain task success, and a full-duplex voice mode. It does not primarily evaluate general desktop or browser control."} {"id":"nemo-voice-agent-evaluation","name":"NVIDIA NeMo Voice Agent Evaluation","primary_focus":"Reproducible live voice-agent harness for EVA and tau2 task domains","audio_input":"yes","audio_output":"yes","multi_turn":"yes","tool_use":"yes","goal_completion":"yes","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"yes","public_data":"yes","code_license":"Apache-2.0","data_license":"MIT for ported EVA and tau2 fixtures; Apache-2.0 for original project material","github_url":"https://github.com/NVIDIA-NeMo/labs-Voice-Agent","huggingface_url":null,"paper_url":null,"last_verified":"2026-09-01","evidence_notes":"The repository ships a live bot-to-bot audio harness with 328 ported EVA and tau2 scenarios, tool execution, stateful outcome checks, resumable runs, and six documented success signals. It is a reproducible implementation layer over existing benchmark domains rather than a new independent task corpus."} {"id":"openbench","name":"OpenBench","primary_focus":"Reproducible ASR, diarization, orchestration, and streaming-transcription benchmarks","audio_input":"yes","audio_output":"no","multi_turn":"no","tool_use":"no","goal_completion":"no","computer_or_browser_action":"no","meeting_or_long_form":"yes","real_time":"yes","public_data":"yes","code_license":"MIT","data_license":"varies","github_url":"https://github.com/argmaxinc/OpenBench","huggingface_url":null,"paper_url":null,"last_verified":"2026-09-01","evidence_notes":"OpenBench measures speech infrastructure and model quality, including streaming and diarization. It is not an end-to-end agent task-completion benchmark."} {"id":"openbenchmarks-voice-agent-latency","name":"OpenBenchmarks Voice Agent Latency","primary_focus":"Caller-perceived time to first agent audio measured from real phone calls","audio_input":"yes","audio_output":"yes","multi_turn":"yes","tool_use":"no","goal_completion":"no","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"yes","public_data":"yes","code_license":"MIT","data_license":"CC-BY-4.0","github_url":"https://github.com/openbenchmarks-labs/voice-agent-latency","huggingface_url":null,"paper_url":null,"last_verified":"2026-09-01","evidence_notes":"The public mirror provides the caller harness, offline analyzer, per-turn artifacts, configuration receipts, recording references and checksums, and a verifier for TTFAB measured from real phone-call audio. It deliberately measures latency only, not answer quality, interruption handling, task completion, or feature breadth."} {"id":"mu-bench","name":"mu-bench","primary_focus":"Multilingual customer-service ASR","audio_input":"yes","audio_output":"no","multi_turn":"no","tool_use":"no","goal_completion":"no","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"no","public_data":"yes","code_license":"Apache-2.0","data_license":"CC-BY-NC-4.0","github_url":"https://github.com/sierra-research/mu-bench","huggingface_url":"https://huggingface.co/datasets/sierra-research/mu-bench","paper_url":null,"last_verified":"2026-09-01","evidence_notes":"Public, gated audio covers 4,270 real customer-service utterances across five locales for ASR provider evaluation. The repository specifies Apache-2.0 for code and CC BY-NC 4.0 for data; this is not a spoken-agent action benchmark."} {"id":"elitr-bench","name":"ELITR-Bench","primary_focus":"Long-context LLM evaluation on meeting transcripts","audio_input":"no","audio_output":"no","multi_turn":"yes","tool_use":"no","goal_completion":"partial","computer_or_browser_action":"no","meeting_or_long_form":"yes","real_time":"no","public_data":"yes","code_license":"BSD-3-Clause main code; Apache-2.0 notices for specified third-party files","data_license":"CC-BY-4.0","github_url":"https://github.com/utter-project/ELITR-Bench","huggingface_url":null,"paper_url":"https://arxiv.org/abs/2403.20262","last_verified":"2026-09-01","evidence_notes":"The benchmark evaluates long-context language-model behavior over meeting transcripts in single- and multi-turn modes. It starts from text transcripts rather than measuring audio capture, ASR, or computer action; the repository publishes separate code and data license files."} {"id":"audio-agent-bench-suite","name":"Audio Agent Bench Suite","primary_focus":"A collection of multi-turn spoken-agent benchmarks","audio_input":"yes","audio_output":"unclear","multi_turn":"yes","tool_use":"yes","goal_completion":"partial","computer_or_browser_action":"no","meeting_or_long_form":"no","real_time":"unclear","public_data":"partial","code_license":"not stated","data_license":"CC-BY-4.0","github_url":null,"huggingface_url":"https://huggingface.co/datasets/arcada-labs/audio-agent-bench-suite","paper_url":null,"last_verified":"2026-09-01","evidence_notes":"The public card links six spoken benchmark datasets spanning instruction following, knowledge grounding, function calls, memory, and state. The suite repository itself exposes the card but no substantive data files; its documented scoring compares responses with golden text and does not clearly establish spoken-output evaluation."}