Keywords:
- Accent Identification
- Accented speech
- Accentual mismatch
- Acoustic model adaptation
- acoustic modeling
- Ad hoc microphone array calibration
- Ad-hoc microphone calibration
- adaptive budget allocation
- adaptive layer norm
- adaptive training
- adaptive TTS
- Adequacy of diffuseness
- Afrikaans
- All pass warp
- ASR
- audio processing
- aurora
- Automatic prosodic event detection
- Automatic Speech Recognition
- Bayesian transfer learning
- Benchmarking
- benchmarks
- bilingual speakers
- Blizzard Challenge
- Broadband beam-pattern
- broadcast news
- Cadzow algorithm
- catastrophic forgetting
- cepstral normalisation
- cochlear model
- cochlear models
- Code-Switching
- conditional layer normalization
- Confidence Measure (CM)
- connectionist temporal classification (ctc)
- constrained structural maximum a posteriori linear regression
- continuous F0 coding
- Conversational technologies
- crosslingual adaptation
- Deep learning for speech
- deep MLPs
- deep neural networks
- Delay-and-sum beamformer
- dialectal lexicon
- Diffuse field coherence model
- Diffuse noise coherence
- Diffuse sound coherence model
- diffusion model
- diffusion transformer
- Digital IIR Filters
- Digital IIR Filters
- direction of arrival
- Directivity
- Distant speech recognition
- Distributed source localization.
- dnn
- dnn-based speech recognition
- domain adaptation
- duration
- Emotion Recognition
- emotional speech synthesis
- emotional TTS
- emphasis
- end-to-end
- end-to-end architectures
- energy
- Environmental mismatch
- Euclidean distance matrix
- fast adaptation
- fast training
- filterbanks
- French accents
- French Regional Accents
- French TTS
- Fujisaki Model
- gamma-tone filter
- Generalized Trust Region Subproblem (GTRS).
- German
- German language
- GMM Modelling
- hidden Markov models
- HMM-based speech synthesis
- HSMM explicit duration modelling
- hybrid system
- i-vectors
- Image Model
- importance score
- intonation
- KL-HMM
- Kullback-Leibler divergence
- Laplace approximation
- Lexicon
- low bit rate speech coding
- low-rank adaptation
- LVCSR
- Matrix completion
- modelling
- multi-dialect
- multilayer perceptron
- Multilingual
- multilingual acoustic modeling
- multilingual ASR
- multilingual speech recognition
- Multimodal interaction
- nearest neighbour rule of classification.
- neural network features
- neural networks
- NLP
- Noise Robustness
- open vocabulary
- open-vocabulary
- Out-Of-Language (OOL) detection
- Overlapping Speech
- parameter-efficient fine-tuning
- parametric speech synthesis
- Parametric vocoding
- pattern matching
- phone duration modelling
- Phonological features
- phonological posteriors
- phonology
- pitch analysis
- pitch model
- Pitch modelling
- pitch target approximation
- pitch target realisation
- Posterior features
- pretrained language model
- probabilistic amplitude demodulation
- prosody
- Prosody Modelling
- punctuation
- real-time audio processing
- recurrent neural network
- reliability estimation
- Reverberant enclosure
- Robust microphone placement
- S-stress
- Saliency Mapping
- self-supervision
- Semi-supervised training
- Semidefinite programming
- sentence boundary prediction
- SGMM adaptation
- SincNet
- Single-channel source localization
- SNR spectrum
- Source localization
- Sparse Component Analysis
- speaker adaptation
- spectral amplitude modulation phase hierarchy
- speech coding
- speech corpus
- speech meta-data
- speech prosody
- speech recognition
- speech synthesis
- Speech Translation
- speech-to-speech translation
- spiking neural networks
- Spoken Language Understanding
- Spoken Term Detection (STD)
- Statistical parametric speech synthesis
- Subs-ace Gaussian Mixture Models
- subword segmentation
- Subword unit
- Superdirective beamformer
- Support Vector Regression
- SVM
- Swiss German
- Swiss prosody
- Swisscom
- synchronisation
- Tandem
- temporal alignment
- text-to-speech
- time synchronisation
- time synchronization
- time-frequency analysis
- trainable filterbanks
- TTS
- TV Box
- Under-resourced data
- under-resourced languages
- under-resourced speech recognition
- unit selection
- universal phoneme set
- VAE
- variational inference
- Very low bit rate speech coding
- vocal tract length normalization
- voice assistant
- VTLN
- Wav2vec
- word emphasis
- zero-shot speaker adaptation
Publications of Philip N. Garner sorted by first author
A
A t-distribution based operator for enhancing out of distribution robustness of neural network classifiers, and , in: IEEE Signal Processing Letters, 27:1070-1074, 2020 |
[DOI] |
Sparse Component Analysis for Speech Recognition in Multi-Speaker Environment, , and , in: Proceedings of Interspeech, Makuhari, Japan, 2010 |
|
B
Exploring neural oscillations during speech perception via surrogate gradient spiking neural networks, and , in: Frontiers in Neuroscience, 18(1449181), 2024 |
[DOI] |
A surrogate gradient spiking baseline for speech command recognition, and , in: Frontiers in Neuroscience, 2022 |
[DOI] [URL] |
Bayesian Recurrent Units and the Forward Backward Algorithm, and , in: Proc. Interspeech 2022, pages 4137-4141, 2022 |
[DOI] |
A Bayesian Interpretation of the Light Gated Recurrent Unit, and , in: Proceedings IEEE International Conference on Acoustics, Speech and Signal Processing, 2021 |
[DOI] |
Current trends in multilingual speech processing, , , , , , , , and , in: Sadhana, 36(5):885–915, 2011 |
[DOI] [URL] |
Idiap Scientific Report 2022, , , , , , , , , , , , , , , , , and , Idiap-RR-05-2023 |
|
C
Sound Pattern Matching for Automatic Prosodic Event Detection, , , , and , Idiap-RR-03-2016 |
|
Sound Pattern Matching for Automatic Prosodic Event Detection, , , , and , in: Interspeech, San Francisco, USA, 2016 |
|
PhonVoc: A Phonetic and Phonological Vocoding Toolkit, and , in: Interspeech, San Francisco, USA, 2016 |
|
Incremental Syllable-Context Phonetic Vocoding, , , , and , Idiap-RR-05-2015 |
|
Incremental Syllable-Context Phonetic Vocoding, , , , and , in: IEEE/ACM TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING, 23(6), 2015 |
[URL] |
Progress report of a project in very low bit-rate speech coding, , and , Idiap-RR-08-2012 |
|
Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech Coding, , , and , Idiap-RR-11-2016 |
|
Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech Coding, , , and , in: IEEE/ACM Trans. on Audio, Speech and Language Processing, 2016 |
|
Stress and Accent Transmission In HMM-Based Syllable-Context Very Low Bit Rate Speech Coding, , , and , in: Interspeech, 2014 |
|
Stress and Accent Transmission In HMM-Based Syllable-Context Very Low Bit Rate Speech Coding, , , and , Idiap-RR-10-2014 |
|
ON THE (UN)IMPORTANCE OF THE CONTEXTUAL FACTORS IN HMM-BASED SPEECH SYNTHESIS AND CODING, , and , Idiap-RR-06-2013 |
|
On the (Un)importance of the Contextual Factors In HMM-Based Speech Synthesis, , and , in: Proceedings of the IEEE Intl. Conference on Acoustics, Speech and Signal Processing (ICASSP), Vancouver, Canada, pages 8140 - 8143, 2013 |
|
Syllable-based Pitch Encoding for Low Bit Rate Speech Coding with Recognition/Synthesis Architecture, , and , in: Proc. of Interspeech 2013, Lyon, France, 2013 |
|
Syllable-based Pitch Encoding for Low Bit Rate Speech Coding with Recognition/Synthesis Architecture, , and , Idiap-RR-24-2013 |
|
Phonological vocoding using artificial neural networks, , and , Idiap-RR-04-2015 |
|
Phonological Vocoding Using Artificial Neural Networks, , and , in: IEEE 40th International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, Australia, pages 4844-4848, IEEE, 2015 |
[DOI] |
A Bayesian Interpretation of Adaptive Low-Rank Adaptation, and , in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2025 |
Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic Forgetting, and , in: IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024 |
[DOI] |
Diffusion Transformer for Adaptive Text-to-Speech, and , in: Proc. 12th ISCA Speech Synthesis Workshop (SSW 12), 2023 |
[DOI] |
The Idiap Speech Synthesis System for the Blizzard Challenge 2023, , , and , in: Proc. 18th Blizzard Challenge Workshop, 2023 |
[DOI] |
Training a Filter-Based Model of the Cochlea in the Context of Pre-Trained Acoustic Models, and , in: Acoustics, 6:470 - 488, 2024 |
[DOI] |
Low-Level Physiological Implications of End-to-End Learning for Speech Recognition, and , in: Proc. Interspeech 2022, pages 749--753, 2022 |
[DOI] |
D
COMBINING CEPSTRAL NORMALIZATION AND COCHLEAR IMPLANT-LIKE SPEECH PROCESSING FOR MICROPHONE ARRAY-BASED SPEECH RECOGNITION, , and , in: Proceedings of the IEEE Workshop on Spoken Language Technology, 2012 |
|
IMPROVING MICROPHONE ARRAY SPEECH RECOGNITION WITH COCHLEAR IMPLANT-LIKE SPECTRALLY REDUCED SPEECH, , and , Idiap-RR-40-2011 |
|
G
Modeling Unvoiced Sounds In Statistical Parametric Speech Synthesis with a Continuous Vocoder, , , and , in: Proc. of EUSIPCO, Budapest, Hungary, 2016 |
|
Combining the SNR Spectrum with a Cochlear Model, , Idiap-RR-14-2018 |
|