Keywords:
- Accent Identification
- Accented speech
- Accentual mismatch
- Acoustic model adaptation
- acoustic modeling
- Ad hoc microphone array calibration
- Ad-hoc microphone calibration
- adaptive budget allocation
- adaptive layer norm
- adaptive training
- adaptive TTS
- Adequacy of diffuseness
- Afrikaans
- All pass warp
- ASR
- audio processing
- aurora
- Automatic prosodic event detection
- Automatic Speech Recognition
- Bayesian transfer learning
- Benchmarking
- benchmarks
- bilingual speakers
- Blizzard Challenge
- Broadband beam-pattern
- broadcast news
- Cadzow algorithm
- catastrophic forgetting
- cepstral normalisation
- cochlear model
- cochlear models
- Code-Switching
- conditional layer normalization
- Confidence Measure (CM)
- connectionist temporal classification (ctc)
- constrained structural maximum a posteriori linear regression
- continuous F0 coding
- Conversational technologies
- crosslingual adaptation
- Deep learning for speech
- deep MLPs
- deep neural networks
- Delay-and-sum beamformer
- dialectal lexicon
- Diffuse field coherence model
- Diffuse noise coherence
- Diffuse sound coherence model
- diffusion model
- diffusion transformer
- Digital IIR Filters
- Digital IIR Filters
- direction of arrival
- Directivity
- Distant speech recognition
- Distributed source localization.
- dnn
- dnn-based speech recognition
- domain adaptation
- duration
- Emotion Recognition
- emotional speech synthesis
- emotional TTS
- emphasis
- end-to-end
- end-to-end architectures
- energy
- Environmental mismatch
- Euclidean distance matrix
- fast adaptation
- fast training
- filterbanks
- French accents
- French Regional Accents
- French TTS
- Fujisaki Model
- gamma-tone filter
- Generalized Trust Region Subproblem (GTRS).
- German
- German language
- GMM Modelling
- hidden Markov models
- HMM-based speech synthesis
- HSMM explicit duration modelling
- hybrid system
- i-vectors
- Image Model
- importance score
- intonation
- KL-HMM
- Kullback-Leibler divergence
- Laplace approximation
- Lexicon
- low bit rate speech coding
- low-rank adaptation
- LVCSR
- Matrix completion
- modelling
- multi-dialect
- multilayer perceptron
- Multilingual
- multilingual acoustic modeling
- multilingual ASR
- multilingual speech recognition
- Multimodal interaction
- nearest neighbour rule of classification.
- neural network features
- neural networks
- NLP
- Noise Robustness
- open vocabulary
- open-vocabulary
- Out-Of-Language (OOL) detection
- Overlapping Speech
- parameter-efficient fine-tuning
- parametric speech synthesis
- Parametric vocoding
- pattern matching
- phone duration modelling
- Phonological features
- phonological posteriors
- phonology
- pitch analysis
- pitch model
- Pitch modelling
- pitch target approximation
- pitch target realisation
- Posterior features
- pretrained language model
- probabilistic amplitude demodulation
- prosody
- Prosody Modelling
- punctuation
- real-time audio processing
- recurrent neural network
- reliability estimation
- Reverberant enclosure
- Robust microphone placement
- S-stress
- Saliency Mapping
- self-supervision
- Semi-supervised training
- Semidefinite programming
- sentence boundary prediction
- SGMM adaptation
- SincNet
- Single-channel source localization
- SNR spectrum
- Source localization
- Sparse Component Analysis
- speaker adaptation
- spectral amplitude modulation phase hierarchy
- speech coding
- speech corpus
- speech meta-data
- speech prosody
- speech recognition
- speech synthesis
- Speech Translation
- speech-to-speech translation
- spiking neural networks
- Spoken Language Understanding
- Spoken Term Detection (STD)
- Statistical parametric speech synthesis
- Subs-ace Gaussian Mixture Models
- subword segmentation
- Subword unit
- Superdirective beamformer
- Support Vector Regression
- SVM
- Swiss German
- Swiss prosody
- Swisscom
- synchronisation
- Tandem
- temporal alignment
- text-to-speech
- time synchronisation
- time synchronization
- time-frequency analysis
- trainable filterbanks
- TTS
- TV Box
- Under-resourced data
- under-resourced languages
- under-resourced speech recognition
- unit selection
- universal phoneme set
- VAE
- variational inference
- Very low bit rate speech coding
- vocal tract length normalization
- voice assistant
- VTLN
- Wav2vec
- word emphasis
- zero-shot speaker adaptation
Publications of Philip N. Garner sorted by first author
S
Study of Jacobian Normalization for VTLN, , and , Idiap-RR-25-2010 |
|
VTLN Adaptation for Statistical Speech Synthesis, , , and , in: Proceedings of ICASSP, Dallas, Texas, 2010 |
|
VTLN Adaptation for Statistical Speech Synthesis, , , and , Idiap-RR-41-2009 |
|
VTLN-Based Rapid Cross-Lingual Adaptation for Statistical Parametric Speech Synthesis, , , and , Idiap-RR-12-2012 |
|
Combining Vocal Tract Length Normalization with Hierarchical Linear Transformations, , , and , in: IEEE Journal of Selected Topics in Signal Processing - Special Issue on Statistical Parametric Speech Synthesis, 8(2):262 - 272, 2014 |
[DOI] |
Bias Adaptation for Vocal Tract Length Normalization, , , and , Idiap-RR-12-2013 |
|
Combining Vocal Tract Length Normalization with Linear Transformations in a Bayesian Framework, , , and , Idiap-RR-11-2012 |
|
COMBINING VOCAL TRACT LENGTH NORMALIZATION WITH HIERARCHIAL LINEAR TRANSFORMATIONS, , , and , in: Proceedings in International conference on Speech and Signal processing, Kyoto, Japan, pages 4493-4496, IEEE SPS (ICASSP), 2012 |
|
Investigating a neural all pass warp in modern TTS applications, and , in: Speech Communication, 138:26--37, 2022 |
[DOI] |
Improving Emotional TTS with an Emotion Intensity Input from Unsupervised Extraction, and , in: 11th ISCA Speech Synthesis Workshop, 2021 |
[URL] |
Neural VTLN for Speaker Adaptation in TTS, and , in: Proc. 10th ISCA Speech Synthesis Workshop, ISCA, Vienna, Austria, pages 6, 2019 |
[DOI] |
A Neural Model to Predict Parameters for a Generalized Command Response Model of Intonation, and , Idiap-RR-10-2018 |
|
A Neural Model to Predict Parameters for a Generalized Command Response Model of Intonation, and , in: Proc. Interspeech 2018, pages 3147-3151, 2018 |
[DOI] |
A Neural Model to Predict Parameters for a Generalized Command Response Model of Intonation, and , in: MLSLP-18 Proceedings, Hyderabad, 2018 |
[URL] |
Design of a Speech Corpus for Research on Cross-Lingual Prosody Transfer, , , , , , , , , , , and , in: Lecture Notes in Artificial Intelligence: 18th International Conference, SPECOM 2016, Budapest, Hungary, pages 199--206, 2016 |
|
Intuitive Recipes for Uncertainty Decoding with SNR Features for Noise Robust ASR, and , Idiap-RR-23-2011 |
|
Automatic Speech Indexing System of Bilingual Video Parliament Interventions, , , , , and , Idiap-RR-25-2013 |
|
The SP2 SCOPES Project on Speech Prosody, , , , , , , , and , in: DOGS2014 - Digital speech and image processing, 2014 |
|
Evaluating Intra- and Crosslingual Adaptation for Non-native Speech Recognition in a Bilingual Environment, and , in: Proceedings of the 4th IEEE International Conference on Cognitive Infocommunications, IEEE, Budapest, Hungary, pages 357-361, 2013 |
|
T
Ad-Hoc Microphone Array Calibration from Partial Distance Measurements, , , and , in: Proceeding of 4th Joint Workshop on Hands-free Speech Communication and Microphone Arrays, 2014 |
Ad-Hoc Microphone Array Calibration from Partial Distance Measurements, , , and , in: Proceedings of the 4th Joint Workshop on Hands-free speech communication and Microphone Arrays, Villers-les-Nancy, pages 1 - 5, IEEE, 2014 |
[DOI] |
Spatial Sound Localization via Multipath Euclidean Distance Matrix Recovery, , , , and , in: IEEE Journal of Selected Topics in Signal Processing, 9(5):802-814, 2015 |
|
Enhanced Diffuse Field Model for Ad Hoc Microphone Array Calibration, , and , in: Signal Processing, 101:242-255, 2014 |
|
Microphone Array Beampattern Characterization for Hands-free Speech Applications, , and , in: IEEE 7th Sensor Array and Multichannel Signal Processing Workshop(SAM), Hoboken, NJ, USA, pages 473-476, 2012 |
|
BROADBAND BEAMPATTERN FOR MULTI-CHANNEL SPEECH ACQUISITION AND DISTANT SPEECH RECOGNITION, , and , Idiap-RR-39-2011 |
|
AN INTEGRATED FRAMEWORK FOR MULTI-CHANNEL MULTI-SOURCE LOCALIZATION AND VOICE ACTIVITY DETECTION, , , , and , Idiap-RR-16-2011 |
|
An Integrated Framework for Multi-Channel Multi-Source Localization and Voice Activity Detection, , , , and , in: The Third Joint Workshop on Hands-free Speech Communication and Microphone Arrays, 2011 |
|
Robust Microphone Placement for Source Localization from Noisy Distance Measurements, , , , and , in: IEEE 40th International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2579-2583, IEEE, 2015 |
[DOI] |
Euclidean Distance Matrix Completion for Ad-hoc Microphone Array Calibration, , , and , in: Proceedings IEEE International Conference On Digital Signal Processing, 2013 |
|
Ad Hoc Microphone Array Calibration: Euclidean Distance Matrix Completion Algorithm and Theoretical Guarantees, , , , and , in: Signal Processing, 107:123–140, 2015 |
[DOI] |
Self-attention for Speech Emotion Recognition, , and , in: Proc. Interspeech 2019, 2019 |
[DOI] |
Phonological mappings for English, French, German and Portuguese, and , Idiap-Com-02-2018 |
|
AN INVESTIGATION OF MULTILINGUAL ASR USING END-TO-END LF-MMI, , and , in: International Conference on Acoustics, Speech and Signal Processing, 2019 |
|
Multilingual Training and Cross-lingual Adaptation on CTC-based Acoustic Model, , and , Idiap-RR-01-2018 |
|
Fast Language Adaptation Using Phonological Information, , and , in: Proceedings of Interspeech 2018, Hyderabad, INDIA, pages 2459-2463, 2018 |
[DOI] |
Cross-lingual Adaptation of a CTC-based multilingual Acoustic Model, , and , in: Speech Communication, 104:39-46, 2018 |
[DOI] |
An Investigation of Deep Neural Networks for Multilingual Speech Recognition Training and Adaptation, , and , in: Proc. of Interspeech, 2017 |
|