• Services
    LLM
    AI & ML
    Digital Healthcare
    Data Science
    DevOps
  • Products
    Jackalope
    EyeAI
  • Industries
    Healthcare
    Agriculture
    EdTech / LMS
    Retail / E-commerce
    Manufacturing
  • Resources
    Blog
    Case Studies
    Expert Guides
  • Company
    About us
    Careers
  • Contact us
logo
Services
LLMAI & MLDigital HealthcareData ScienceDevOps
Industries
HealthcareAgricultureEdTech / LMSRetail / E-commerceManufacturing
Case StudiesAbout UsBlogCareers
Our contacts
+380(66)54-32-579
sales@sciforce.tech

Get monthly digest of innovations

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Social Media:
Privacy Policy © 2026 Sciforce
5.0
Interspeech

Hot Topics at Interspeech 2024: The Latest in Technology of Spoken Language Processing

Published: October 1, 2024
# AI / ML
# Speech Processing

What’s Hot at Interspeech 2024

Our team recently attended the 25th Interspeech Conference, held from September 1st to 5th on Kos Island, Greece. This year’s theme, "Speech and Beyond," highlighted new developments in speech technology, focusing on areas like healthcare diagnostics, virtual assistants, and even animal sound recognition. It was a great opportunity for experts worldwide to share their work and discuss the latest trends. Here are some of the key topics and insights we gathered from the event.

Interspeech.jpg

Using Large Language Models for ASR and SLU

A major topic at the conference was the use of Large Language Models (LLMs) to improve Automatic Speech Recognition (ASR) and Spoken Language Understanding (SLU) systems. Unlike traditional ASR models, LLMs are trained on large amounts of text data, which helps them better understand the context of spoken words. This makes them particularly useful for better recognition on domains with missing or scarce acoustic training data.

The research team (Jinlong Xue, Yayue Deng, Yicheng Han, Yingming Gao, Ya Li) showed how they are using LLMs to fix errors in ASR outputs, while others are exploring new methods for real-time speech recognition. There is also a growing trend of adding more text data to the training process for acoustic models, which improves their ability to recognize less common words and phrases. This approach is gaining popularity as it leads to better overall performance of ASR systems.

To learn more about how to create your own large language model, check out our step-by-step guide.

Improving the Whisper Model

The Whisper model and its updated versions were one of key topics at the conference. The research is focused on improving Whisper’s performance, such as making the text alignment more accurate, reducing errors, and better detecting pauses in speech. One example is Crisper Whisper, – a modified version originally designed for medical diagnostics. It offers more precise timing and fewer mistakes, making it a strong alternative to WhisperX for various uses, like legal transcription where high accuracy is needed.

The conference also highlighted several Whisper-like models using open datasets, which could make these advanced tools available for a wider range of applications. This growing ecosystem of Whisper-based models shows promise for improving accessibility and adapting to different industries and languages.

New Ways to Detect Fake Audio

As synthetic voice technology improves, it’s becoming more important to detect fake audio, known as deep fakes. The paper highlighted several new methods for spotting these manipulations. Some researchers are using advanced models to identify small differences between real and fake voices, even when the fake is very convincing.

These developments are crucial for keeping voice data secure in areas like security, media, and entertainment, where it’s important to trust that audio is genuine.

Better Speech Recognition for People with Speech Disorders

In the healthcare field, a major update was the release of a new dataset for dysarthric speech by Mark Hasegawa-Johnson, the creator of the UASpeech corpus. This dataset is designed to help improve speech recognition for people with severe speech impairments and will be used in a research challenge this November to test new models and approaches.

Google continues the development of its project, previously known as Euphonia, which aims to enhance speech recognition for individuals with speech disorders. They presented new techniques that make their models better at understanding and transcribing slurred or irregular speech, benefiting users with conditions like cerebral palsy or ALS.

These advancements are essential for developing technology that can accurately recognize and respond to the unique speech patterns of people with dysarthria — an area where personalized adaptation has shown dramatic accuracy gains.

Speech Technology for E-Learning

ASR (Automatic Speech Recognition) and speech technology in e-learning were well-covered topics at the conference. Many of the techniques presented were similar to the ones we developed five years ago. For example, using phonological features for mispronunciation detection. The paper is describing using a wav2vec2 model with modified CTC loss, while we used a transformer model for similar tasks as well.

What’s Next for Speech Technology

The conference offered valuable insights into speech technology's latest trends and future possibilities. We’re looking forward to using these ideas in our current projects and working with the research community to explore new possibilities.

We plan to use LLMs to improve ASR accuracy, adopt Whisper model enhancements for specialized transcriptions, and refine our speech recognition capabilities for individuals with disorders. For e-learning, we're exploring real-time feedback solutions for better pronunciation training. Beyond research, these ASR advances are already powering real-world deployments like voice-driven ordering in drive-thru chains.

Stay tuned for more in-depth reports on the topics discussed at the event. If you have any questions or would like to talk about any of these findings, feel free to reach out!

RELATED BLOG ARTICLES

View all Articles
OHDSI Europe Symposium 2026From OMOP Workflows to Living Evidence: SciForce at OHDSI Europe Symposium 2026

This April, Polina Talapova and Mariia Pahur represented SciForce at the 7th European OHDSI Symposium in Rotterdam – three vivid days of workshops, poster sessions, MindMeetsMachines mapping competition and an oral presentation aboard the SS Rotterdam, a retired ocean liner moored on the Maas river. The symposium's theme was Continuous Collaboration for Living Evidence Generation. The word "living" matters here. Traditional evidence-generation projects are often designed as discrete studies. A

# Healthcare
# AI / ML
# Data Science
# LLM
Telehealth Platform ArchitectureTelehealth Platform Architecture: Building Secure, Scalable Virtual Care Systems

Building a telehealth platform at clinical scale means solving for hospital network restrictions, HIPAA compliance and auditability, and the data load of continuous remote monitoring – and the architecture decisions that determine whether it holds up are mostly made in the first few sprints. The engineering debt from early decisions starts showing up at scale: video sessions dropping when hospital firewalls, restrictive egress policies, or network address translation prevent a direct media path;

# Healthcare
# AI / ML
# Data Science
Improving Diagnostic Accuracy and WorkflowAI in Medical Imaging: From Diagnostic Accuracy to Clinically Usable Workflow

A radiologist on a standard hospital shift may read dozens to well over a hundred imaging studies, depending on subspecialty, setting, shift structure, and case complexity. Each one is a search for something that might be subtle, easy to miss, or buried in noise. At that volume, non-trivial discrepancy or error rate is a known risk in radiology practice, especially under high workload and time pressure. Radiologists are working through growing imaging volumes with a workforce that has never full

# Healthcare
# AI / ML
# Computer Vision
# Data Science
Sustainable AI: Strategies for Managing Compute Costs and Energy EfficiencySustainable AI: Strategies for Managing Compute Costs and Energy Efficiency

In 2025, the world’s data centers consumed 485 terawatt-hour of energy, with AI-related demand growing at 50%. By 2030, the consumption is expected to reach 950 TWh – twice as much as today, and equals approximately the entire electricity consumption of Japan. Goldman Sachs forecasts that about 60% of new demand will be met by burning fossil fuels, increasing global carbon emissions to 220 million tons. And as the chart below shows, the emissions cost escalates sharply with each new generation o

# AI / ML
# Data Science