IEEE Transactions on Speech and Audio Processing
What Is IEEE Transactions on Speech and Audio Processing?
IEEE Transactions on Speech and Audio Processing was a peer-reviewed journal that published research on the analysis, synthesis, coding, recognition, and enhancement of speech and audio signals from 1993 to 2005. A publication of the IEEE Signal Processing Society, it served as the primary archival venue for a disciplinary community that required more focused coverage than the broader IEEE Transactions on Acoustics, Speech, and Signal Processing (ASSP) had provided since 1974. When the ASSP Transactions was reorganized in 1992, the signal processing side became IEEE Transactions on Signal Processing, and the speech and audio community received its own dedicated journal.
The journal drew on digital signal processing, acoustics, linguistics, and psychoacoustics to address the full range of problems arising when humans communicate through sound. Its scope ran from the physics of vocal tract production through the perceptual mechanisms of hearing and into the computational and statistical methods used to process, code, and understand spoken and musical audio.
Speech Analysis, Synthesis, and Coding
A central area of the journal covered methods for analyzing and parametrically representing speech. Linear predictive coding (LPC) and its variants provided compact parametric models of the vocal tract filter and drove much of the speech coding work in early volumes. Code-excited linear prediction (CELP), the basis for most narrow-band voice codecs standardized through the ITU-T and later the IETF, appeared in the journal as both theoretical work and subjective evaluation studies.
Speech synthesis, concerned with generating intelligible and natural-sounding speech from text or phonemic input, was covered through work on formant synthesis, concatenative synthesis using stored speech units, and prosody modeling. Evaluation methods for synthesized speech quality, including both objective measures and perceptual test protocols, appeared alongside the synthesis algorithms themselves. These developments fed directly into the voice communication standards developed by bodies such as the ITU-T, whose G-series codec recommendations drew on research published in this and related venues.
Speech Recognition and Language Modeling
The journal tracked the rapid evolution of automatic speech recognition (ASR) from hidden Markov model (HMM) based acoustic modeling through the introduction of neural network approaches. Feature extraction methods, including mel-frequency cepstral coefficients (MFCCs) and perceptual linear prediction (PLP), were subjects of comparative studies. The integration of language models with acoustic models, decoding algorithms including Viterbi search and its extensions, and speaker adaptation methods all appeared in the Transactions.
Speaker recognition, the task of identifying or verifying a speaker's identity from a voice sample, received dedicated coverage as a distinct problem from speech recognition. Work on Gaussian mixture model (GMM) based speaker verification, i-vector representations, and robustness to channel and environmental variation appeared as the technology found applications in telephone-based authentication. The IEEE Signal Processing Society maintained the publication as part of a portfolio that also included the Signal Processing Letters, creating a complementary set of venues for short and full-length contributions.
Transition to Audio, Speech, and Language Processing
In 2006, the journal was succeeded by IEEE Transactions on Audio, Speech, and Language Processing, which added explicit coverage of natural language processing and music signal analysis. This renaming reflected the convergence of the speech and audio community with language processing research, particularly as statistical methods became central to both automatic speech recognition and machine translation. The archived volumes of the original Transactions on Speech and Audio Processing, accessible through IEEE Xplore, document the development of the acoustic modeling and signal processing foundations that later deep learning-based systems were built upon.
Applications
IEEE Transactions on Speech and Audio Processing published research with applications in:
- Cellular and VoIP voice coding for efficient telecommunications transmission
- Hands-free voice communication and acoustic echo cancellation
- Voice-controlled interfaces and interactive voice response systems
- Hearing aid signal processing and auditory prosthetics
- Music analysis, retrieval, and automatic transcription