AI Sound Effect Generation: A Narrative Review of Generative Models
A recent academic examination has investigated how various input methods—such as text, images, audio, and their combinations—influence the output quality of sound effects generated by AI systems. This review encompasses findings from 30 scholarly articles retrieved from platforms like Google Scholar, IEEE Xplore, and the ACM Digital Library, reflecting significant developments in AI generative techniques over the last five years. While the review signals noteworthy progress, it does not disclose specifics about individual models or their evaluation metrics. The paper is accessible on arXiv with the identifier 2608.03742 and forms part of a broader scholarly chapter, underscoring rising interest in AI audio applications.
Key facts
- The review analyzes 30 peer-reviewed articles from Google Scholar, IEEE Xplore, and the ACM Digital Library.
- It focuses on how input modalities (text, visual, audio, multimodal) affect generated audio quality, controllability, and contextual relevance.
- The review explores the evolution of AI generative models over the past five years.
- Multiple models achieved state-of-the-art performance, according to the results.
- The paper is available on arXiv with identifier 2608.03742.
- The announcement type is 'cross'.
- The review is part of a chapter, indicating it may be published in a larger work.
- AI-driven audio generative models are rapidly growing in popularity.
Entities
Institutions
- Google Scholar
- IEEE Xplore
- ACM Digital Library
- arXiv