ISCA - International Speech Communication Association
ISCA Archive
Back to list of job offers
We are recruiting a fully funded, three-year PhD candidate for a project on Privacy-Preserving Speech Understanding with Multimodal Signals for Clinical Applications, hosted by the MULTISPEECH team at LORIA (Université de Lorraine, Inria and CNRS), in Villers-lès-Nancy, France.
The project is funded by the AI Grand Est ENACT research chair and will investigate methods to protect speaker identity and sensitive content while preserving the semantic and diagnostic information needed for clinical speech understanding. The work will combine speech and text, with the possibility of incorporating an additional modality such as physiological signals, medical imaging or electronic health-record metadata. It will involve privacy-preserving speech processing, speech/audio foundation models, multimodal machine learning, and evaluation of privacy–utility trade-offs.
We welcome candidates with a Master’s or engineering degree in computer science, AI, signal or speech processing, applied mathematics, data science, or a related area. Strong Python and machine-learning/deep-learning skills are expected. Experience in speech processing, NLP, privacy-preserving ML, multimodal learning, or AI for healthcare is particularly welcome.
Applicants should email a CV, motivation letter, degree transcripts, and, if available, two recommendation letters or the contact details of two referees.
For the full position description and application details, please see: https://sites.google.com/view/natalia-tomashenko/recruitment-phd-positions Contact: Natalia Tomashenko, natalia.tomashenko@inria.fr
Job at Auphonic: Audio Machine/Deep Learning & Signal Processing Engineer (Graz/Austria) Hallo! Auphonic is looking for a Machine/Deep Learning & Signal Processing Engineer in Graz, Austria: "Work on real-world audio AI — from classic DSP to large deep learning models — using our large-scale datasets and GPU infrastructure." More details: https://auphonic.com/jobs/ml-engineer LG Georg
The Speech and Multimodal Intelligent Information Processing (SMIIP) Lab at the Chinese University of Hong Kong, Shenzhen has multiple open positions for fully funded Ph.D. student and Postdoc Researchers (2 years contract).
Our research interests lie in the areas of intelligent speech processing, embodied audition and dialogue system as well as multimodal behavior signal analysis and interpretation.
Email: mingli369@cuhk.edu.cn
PI information: Prof. Ming Li https://sai.cuhk.edu.cn/en/teacher/258
Lab Website: https://smiip-mli.github.io/
PhD admission information: https://sai.cuhk.edu.cn/en/node/35
Postdoc recruitment information: https://www.cuhk.edu.cn/en/taxonomy/term/50
It is highly recommended to contact Prof. Ming Li through email before the application. Thanks.
We’re looking for a strong product leader who can drive roadmap, prioritization, and product direction — and who’s excited by the intersection of language, community, and AI. This role will help lead Common Voice, Mozilla’s global, community-powered platform helping make AI more inclusive by ensuring speakers of the world’s languages can be represented in training data. We’d especially love to hear from people with:
Nice to have experience:
Follow this link to apply: https://job-boards.greenhouse.io/mozilla/jobs/7813301
At the MSAD groupe, within the LIST3N laboratory at the University of Technologi of Troyes, we are offering a research internship.
The aim of this internship is to explore and analyze a new approach based on collaborative knowledge distillation where teacher and student are trained simultaneously.
Each model mutually enriches the other through a reciprocal influence on their learning processes, going beyond traditional unidirectional transfer.
To validate this concept, the approach will be applied to speech recognition based on Connectionist Temporal Classification (CTC).
A known problem with CTC is alignment divergence: models trained separately on the same data often develop inconsistent temporal alignments.
With collaborative learning, we hypothesize that they will naturally converge towards a unified alignment, thus improving robustness and performance.
Main Tasks
Candidate's profile
Intership modalities
Application
If you wish to be considered for this internship opportunity, please send your CV and cover letter to mohammed_faouzi.benzeghiba@utt.fr
The University of Oldenburg, Germany, is seeking to fill a permanent professorship (salary scale W2) of Medical Physics with focus on Technical and Experimental Audiology. For more information about the position, please visit https://uol.de/en/job/medical-physics-tea-1027. The professorship is part of the Department of Medical Physics and Acoustics (https://uol.de/en/mediphysics-acoustics). Please contact Prof. Dr. Volker Hohmann (Email: volker.hohmann@uni-oldenburg.de) if you have any questions.
The Biofeedback Intervention Technology for Speech Lab at NYU (BITS Lab; PI: Tara McAllister), in collaboration with NYU's Music and Audio Research Lab (MARL), is seeking a full-time postdoctoral researcher with expertise in digital signal processing, audio and acoustics, or machine learning and demonstrated experience working with speech data. The project involves developing real-time acoustic biofeedback software for clinical speech intervention, with a focus on automated classification of sibilant productions as accurate or distorted. The postdoc will compare analytic and ML approaches to sibilant classification and contribute to acoustic visualization tools for a web-based clinical application. Candidates from a clinical speech or acoustic phonetics background are also welcome to apply.
The position will begin in summer or fall 2026 and is fully remote for U.S.-based candidates (Eastern or Central time zones preferred). In compliance with NYC's Pay Transparency Act, the annual base salary range for this position is $65,000–$70,000. New York University considers factors such as, but not limited to, the specific grant funding and the terms of the research grant when extending an offer. To apply, submit a cover letter, CV, and 2–3 references via Interfolio: https://apply.interfolio.com/183906. Deadline: April 30, 2026.
BITS Lab is also recruiting PhD students to begin in fall 2027, with an option to be jointly supervised by faculty from MARL. Interested applicants should contact Dr. McAllister directly.
We offer a 3-year position (starting, 01.07.2026, fulltime) to a speech scientist or engineer within the Transregional Collaborative Research Center (TRR 318) “Constructing Explainability,” which is jointly run by the Universities of Paderborn and Bielefeld. The TRR investigates how algorithmic transparency can be promoted, particularly in the context of black-box models as used in modern artificial intelligence systems. The position can be used for further academic qualification.
Our subproject deals with explanatory strategies for recognizing stress in clinical explanatory situations based on multimodal signals (facial expressions and voice). We are developing tools that help clinical staff to better recognize the presence of stress in neurodiverse and neurotypical populations.
For more details, and a link to the application process, see https://jobs.uni-bielefeld.de/job/view/4874/research-position-m-f-d-in-the-sfb-trr-318-in-the-field-of-phonetics?page_lang=en
The Phonetics group at Lancaster University, UK is looking to appoint a Senior Research Associate (i.e. postdoc) in Machine Learning for Speech Processing. The position is available from 1 July 2026 for 18 months.
The goal is to recover vocal tract movements from the acoustic signal. We are developing ways to integrate physical knowledge into the models, so the inversions are not just accurate but also reveal underlying principles of speech production.
More information
4-year fully PhD funded position, supervised by Prof. Naomi Harte in School of Engineering, Trinity College Dublin, Ireland. Research will explore how multimodal cues used in speech-based interaction can be used to track an active speaker in conversations that go beyond controlled, 2-person scenarios. Full details, including how to apply are here:
https://www.adaptcentre.ie/careers/phd-studentship-speaker-tracking-in-complex-conversations/
© Copyright 2024 - ISCA International Speech Communication Association - All right reserved.