The Agentic Data Company announced Open Yap 1K on September 3, 2026. The corpus contains 1,000 hours of natural English conversations between two people, giving developers material that captures the timing problems speech-to-speech systems face outside perfectly separated prompts.

Each speaker has a separate audio track, and both tracks are aligned to the same timeline. The recordings were captured at 48 kHz in real-world environments. That format preserves when each person speaks, pauses, overlaps or responds while the other person is still talking, rather than reducing the exchange to a finished transcript.

A dataset built around ordinary turn-taking

Open Yap 1K was collected through an app comparable to a phone call. One participant invited someone they knew, and the resulting conversations retained interruptions, overlapping speech, laughter and conversational feedback signals. Those details are relevant to full-duplex models, which need to listen and respond continuously while handling the possibility that both speakers may talk at once.

The corpus contains 1,602 conversations from 239 speakers. Conversations average 37.5 minutes, with a 30-minute median, while some last more than 100 minutes. Across the full collection, overlapping speech accounts for a median 8.3% of expressed speaking time. That measurement gives developers a defined view of how often simultaneous speech appears in this material.

The speaker information comes with a clear boundary. Demographic data was reported by the participants rather than inferred from their voices. For evaluation and analysis, that distinction matters: audio-based demographic guesses are not presented as confirmed attributes.

Commercial use is allowed, but full access is reviewed

The complete corpus is offered free for commercial and research use. It is not immediately unrestricted, however. Access requires an application, a review and a data-use agreement. The agreement allows commercial use, research, model training, evaluation and deployment of models trained on the corpus.

It also sets limits on the recordings themselves. Redistribution and resale are prohibited, as are attempts to identify a speaker or create an identifiable voice clone from the data. The arrangement therefore gives teams broad permission to use the corpus for model work while restricting how the source audio can circulate and be repurposed.

A smaller sample is available through the Hugging Face Hub: 8.9 hours covering 16 conversations and eight speakers under a CC-BY-4.0 license. That sample offers a way to inspect the material before applying for the complete collection, but it does not remove the review and contractual requirements attached to full access.

Audio remains important when the transcript is uncertain

Privacy and content controls are part of the delivery process. Every conversation receives human evaluation and LLM-based filtering for personally identifiable information and content-policy violations before delivery. The transcripts have a separate limitation: Deepgram Nova-3 generates them automatically at word level, and humans have not checked them.

That makes the aligned audio the stronger reference when exact wording, interruptions or timing matter. The transcripts can support work with the corpus, but they should not be treated as a human-verified record of every exchange.

Open Yap 1K expands the material available for building and testing conversational AI around real turn-taking. Its practical value lies in the combination of long conversations, aligned speaker tracks and preserved overlap. The trade-off is equally clear: teams must work within access restrictions and account for transcription errors, speaker privacy and the limits of participant-reported demographic information.

Official sources

Sources and methodology

  1. Official source: huggingface.co Opens an external source