JSALT 2026 Plenary Lectures

Participants attend a plenary lecture during the JSALT 2025 workshop in Brno, Czechia.

The JSALT 2026 Plenary Lecture Series features invited talks from leading researchers in speech, language, and artificial intelligence. Lectures are live-streamed via the CLSP YouTube channel.


Lecture Schedule

This year’s lecture series features internationally recognized researchers discussing advances in speech, language, AI, machine learning, and computational linguistics.

Click a lecture below to view its title, abstract, and speaker biography.

  • Date: Wed, July 1, 2026

    Time: 11:00-12:00 followed by lunch

    Venue: Room WT41/43 and DSAI Lounge, 6225 Smith Avenue

    Transport: @10:15 AM, a charted bus will depart from Mason Hall on the Homewood Campus for Mt Washington

    Speaker: John T. Hale, Department of Cognitive Science, Johns Hopkins University

    Title: Co-operating mechanisms of human sentence comprehension

    Abstract: Cognitive scientists seek a computational theory of the mind, including its language-using abilities.  Among those abilities, this talk considers word-by-word sentence comprehension.  Can comprehension itself be resolved into component mechanisms such as memory retrieval and syntactic disambiguation?  Modeling these mechanisms using ideas from computational linguistics, and re-analyzing freely available magnetoencephalography (MEG) data we find support for a temporal dissociation between these two hypothesized components.  This type of analysis serves to refine a kind of “block diagram” for language in the human mind.

    Bio: Prof. Hale received his PhD in Cognitive Science in 2003. He previously taught at Michigan State University, Cornell University, and the University of Georgia. He was a full time research scientist at Google DeepMind 2017-2018. His research uses computational modeling and neuroimaging to study human language processing.

    Photo:

  •  

    Date: Fri, July 10, 2026

    Time: 10:45 A.M.

    Venue: Room WT40 and DSAI Lounge, 6225 Smith Avenue

    Speaker: Ramani Duraiswami, Professor, Department of Computer Science and UMIACS; University of Maryland, College Park

    Title: Towards Auditory General Intelligence: Building, Benchmarking and Improving Large Audio-Language Models

    Abstract: Perception of audio events, music, and speech plays a fundamental role in how humans interact with the world. Large language models have absorbed vast amounts of knowledge from text, but they currently lag in auditory scene understanding, speech and non-speech communication, and music analysis—all central facets of human intelligence. Building AI systems that truly understand audio the way humans do remains an open grand challenge.

    I will introduce the emerging class of Large Audio-Language Models (LALMs)—systems that connect audio encoders to large language models so they can listen, reason, and respond to queries about what they hear. I will start with the basic architectural recipe for teaching a pretrained language model to process audio, covering key design choices: cross-attention versus token prepending, frozen versus finetuned encoders, and curriculum training strategies. I will then trace the evolution of our group’s Audio Flamingo family of models, which over two years has grown from basic audio captioning to structured music understanding, multi-talker speech in 30-minute recordings, temporal chain-of-thought reasoning, and joint audio-visual comprehension. The family spans COMPA, GAMA, Audio Flamingo 2 and 3, Music Flamingo, AF-Next, and Audio-Visual Flamingo, published across ICLR, EMNLP, ICML, NeurIPS, and ICLR 2024–2026. At the time of their publications, these models surpassed both open-weight and closed-source systems across over 20 benchmarks, and remain among the leading fully open models for audio understanding.

    A recurring theme will be the tight coupling between building models and building benchmarks. Benchmarking has been crucial for language model development, yet such benchmarks for LALMs were initially absent. I will describe our creation of MMAU (ICLR 2025)—the first comprehensive benchmark for audio general intelligence, now widely used for evaluating LALMs—and its successor MMAU-Pro (AAAI 2026; a product of JSALT 2025). 

    Finally, I will close with some related broader work in my group, on building large models for other domains. I will sketch how we are applying this same paradigm to genomics, building multimodal embeddings for DNA sequences; and to scientific computing, where our GAIA framework learns neural operators for solving forward and inverse problems based on PDEs on complex geometries.

    Biography: Ramani Duraiswami is a Professor in Computer Science at the University of Maryland, College Park, with appointments in the Artificial Intelligence Institute at Maryland (AIM), UMIACS, Electrical Engineering, Robotics, Neural and Cognitive Sciences, and the Applied Mathematics and Scientific Computing program. His research spans machine learning, scientific computing, and computational perception. His earlier work on fast multipole methods for acoustic scattering, GPU computing, and real-time spatial audio led to two company spinouts; his lab’s audio technology now powers millions of VR headsets, PCs, and headphones worldwide via CEVA. He has published over 400 archival papers across computer science, acoustics, applied mathematics, and machine learning. Prof. Duraiswami holds a B. Tech. from IIT Bombay and a Ph.D. from The Johns Hopkins University.

  • Date: Wed, July 15, 2026

    Time: 10:45 a.m.

    Venue: Terrace Level Room WT40, 6225 Smith Avenue

    Speaker: Jake Beal, RTX BBN Technologies

    Title: LLMs and the Challenge of Keeping Dangerous DNA Out of Risky Hands

    Abstract: Nucleic acid synthesis is both critical for the bioeconomy and an increasingly pressing security concern due to the potential for accidental or deliberate misuse. One of the key lines of defense against misuse is screening nucleic acid orders for potential “sequences of concern.” The advent of Large Language Models (LLMs) and related AI technologies poses critical challenges to this line of defense, greatly increasing the number of potentially capable actors and the potential diversity of sequences of concern. At the same LLMs also provide opportunities for improving defenses, if we can develop methods that allow us to safely rely on the results produced by LLMs. Finally, the lessons being learned at the intersection of AI and biosecurity may prove useful for other AI-impacted domains as well.

    Bio: Dr. Jacob Beal, an Engineering Fellow at RTX BBN Technologies, is the lead developer of FAST-NA Scanner, an adaptation of signature-based malware scanning to DNA synthesis biosecurity screening, now being used as commercial software in the DNA synthesis industry. Dr. Beal also co-leads the Sequence Biosecurity Risk Consortium (SBRC), an effort to establish international standards for testing biosecurity sequence screening systems, as well as organizing efforts to respond to emergent AI-driven biosecurity threats, representing BBN to the International Gene Synthesis Consortium (IGSC), and leading maintenance of the IGSC’s Regulated Pathogen Database.

    Dr. Beal is also known for his work in other areas of synthetic biology, including development of standards for representation and communication of biological designs and experiments, methods for calibrated flow cytometry, precision analysis and design of genetic regulatory networks, and engineering of biological information processing devices.

  • Date: Wed, July 22, 2026

    Time: 10:45 a.m.

    Venue: Terrace Level Room WT40, 6225 Smith Avenue

    Speaker: Shyam Gollakota, University of Washington

    Title: Super-human and proactive audio AI systems

    Abstract: Imagine yourself in a crowded room with a cacophony of sounds, yet having the ability to focus on specific sounds or remove unwanted ones based on their semantic descriptions. This entails understanding and manipulating an acoustic scene in real-time,  isolating each sound and associating spatial context or semantic meaning with it,  a formidable challenge even for the human brain. In this talk, I will showcase a series of projects that demonstrate how AI can augment humans to tackle tasks that are difficult for the human auditory system alone. 

    I will then go beyond auditory augmentation to proactive intelligent systems: hearables that do not merely respond to commands but anticipate and act on our behalf. I will present our recent work on proactive hearing agents and proactive in-ear AI, as well as full-duplex spoken dialogue systems that integrate audio-visual context to listen, see, and speak simultaneously rather than taking rigid turns; pointing toward a shift from AI that augments perception to AI that acts as an always-available, context-aware conversational partner.

    Bio: Shyam Gollakota is a Washington Research Foundation Endowed Professor at the Paul G. Allen School of Computer Science & Engineering at the University of Washington. His work has been licensed by ResMed Inc. and commercialized through his startup, Sound Life Sciences, which was acquired by Google, and is in use by millions of users. He was also CEO of the startup, where he obtained FDA 510(k) clearance for the technology developed in his lab. His lab also worked closely with the Washington State Department of Agriculture to wirelessly track invasive “murder” hornets, which resulted in the destruction of the first nest in the United States. He is the recipient of the ACM Grace Murray Hopper Award in 2020, a Moore Inventor Fellowship in 2021, and the Infosys Prize in 2024. He was also named to MIT Technology Review’s 35 Innovators Under 35, Popular Science’s “Brilliant 10,” and twice to Forbes’ 30 Under 30 list. His group’s research has earned Best Paper awards at MOBICOM, SIGCOMM, UbiComp, SenSys, NSDI, and CHI; has appeared in interdisciplinary journals including Nature, Nature Electronics, Science Robotics, Nature Biomedical Engineering, Science Translational Medicine, and Nature Communications; and has been named an MIT Technology Review Breakthrough Technology of 2016 and one of Popular Science’s top innovations of 2015. He is an alumnus of MIT (Ph.D., 2013; winner of the ACM Doctoral Dissertation Award) and IIT Madras.

    Screenshot

Participants attend a plenary lecture during the JSALT 2025 workshop in Brno, Czechia.

Center for Language and Speech Processing