Visual Bank Inc. (Minato-ku, Tokyo; CEO: Masayuki Nagai) has announced the release of the 'Japanese Regional Dialect Dialogue Speech Dataset' through its AI training data solution, 'Qlean Dataset,' operated by its subsidiary Amana Images Inc.
### About the Dialect Speech Dataset This dataset is a speech corpus containing regional speech patterns, accents, and vocabularies that are not covered by standard language corpora. It is intended for machine learning tasks such as verifying the generalization performance of ASR models, improving dialect understanding in LLMs, and building region-specific TTS models. Custom recordings and additional dialects are also supported upon request.
### Dataset Overview The dataset features natural, spontaneous two-party conversations between Japanese men and women speaking Osaka and Hiroshima dialects. Unlike scripted readings, these recordings capture natural prosody, sentence-ending expressions, and vocabulary, providing acoustic features close to real-world environments. The speaker information includes gender labels, supporting acoustic model evaluation by attribute and adaptation experiments for multi-speaker models.
- **Data Type:** Audio (2-speaker dialogue format) - **Subject Attributes:** Japanese speakers from various regions (with gender labels) - **Capacity:** 5 hours - **Format:** mp3 / wav - **Audio Rate:** 44.1kHz・48kHz / 16・24bit - **Dialects:** Osaka dialect, Hiroshima dialect, etc. - **Commercial Use:** Allowed
### FAQ - **ASR Development:** Can be used for robustness benchmarking (measuring WER) and dialect adaptation using LoRA or full fine-tuning for models like Whisper and ESPnet. - **LLM Development:** Useful for training dialect-to-standard style conversion models and evaluating context-dependent semantic interpretation tasks. - **TTS Applications:** Suitable for fine-tuning models like VITS and StyleTTS to generate natural dialect speech for regional guide robots or dialogue agents. - **Custom Requests:** Custom collection for specific regions, ages, or situations is available.
### Key Use Cases 1. **ASR Robustness Benchmarking:** Quantitative evaluation of recognition accuracy for dialect speech using WER/CER. 2. **Dialect Adaptation Fine-tuning:** Use for few-shot or LoRA fine-tuning to adapt models to specific regional speech. 3. **LLM Dialect Understanding:** Training and evaluation for sentiment analysis, dialect conversion, and discourse structure analysis. 4. **Region-specific TTS Construction:** Building speech generation engines with natural intonation for local services. 5. **Domain Adaptation for Contact Centers:** Developing custom language models for business environments where dialects are frequently used.
### About Qlean Dataset Qlean Dataset is a solution provided by Amana Images Inc. (a Visual Bank subsidiary) offering legally cleared, commercially available AI training data. It covers various formats including audio, image, video, 3D, and text, enabling AI developers to procure high-quality data without legal risks.
FACT BOX
- Source: PR TIMES
- Category: New Product
- Products / services: Qlean Dataset