Close Menu
    Facebook X (Twitter) Instagram
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Facebook X (Twitter) Instagram
    Fintech Fetch
    • Home
    • Crypto News
      • Bitcoin
      • Ethereum
      • Altcoins
      • Blockchain
      • DeFi
    • AI News
    • Stock News
    • Learn
      • AI for Beginners
      • AI Tips
      • Make Money with AI
    • Reviews
    • Tools
      • Best AI Tools
      • Crypto Market Cap List
      • Stock Market Overview
      • Market Heatmap
    • Contact
    Fintech Fetch
    Home»AI News»Meta AI Releases SAM Audio: A State-of-the-Art Unified Model that Uses Intuitive and Multimodal Prompts for Audio Separation
    Meta AI Releases SAM Audio: A State-of-the-Art Unified Model that Uses Intuitive and Multimodal Prompts for Audio Separation
    AI News

    Meta AI Releases SAM Audio: A State-of-the-Art Unified Model that Uses Intuitive and Multimodal Prompts for Audio Separation

    December 17, 20253 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email
    murf

    Meta has released SAM Audio, a prompt-driven audio separation model that targets a common editing bottleneck, isolating one sound from a real-world mix without building a custom model per sound class. Meta released 3 main sizes: sam-audio-small, sam-audio-base, and sam-audio-large. The model is available to download and experiment with in the Segment Anything Playground.

    Architecture

    SAM Audio uses separate encoders for each conditioning signal—an audio encoder for the mixture, a text encoder for the natural language description, a span encoder for time anchors, and a visual encoder that consumes a visual prompt derived from video plus an object mask. The encoded streams are concatenated into time-aligned features, which are then processed by a diffusion transformer that applies self-attention over the time-aligned representation and cross-attention to the textual feature. A DACVAE decoder reconstructs waveforms and emits two outputs: target audio and residual audio.

    What SAM Audio does, and what ‘segment’ means here?

    SAM Audio takes an input recording that contains multiple overlapping sources, like speech plus traffic plus music, and separates a target source based on a prompt. In the public inference API, the model produces two outputs: result.target (the isolated sound) and result.residual (everything else).

    This target-residual interface maps directly to editor operations. For instance, to remove a dog bark from a podcast track, treat the bark as the target and keep only the residual. Conversely, if you want to extract a guitar part from a concert clip, you keep the target waveform instead. Meta uses these examples to illustrate the model’s potential.

    The 3 prompt types Meta is shipping

    Meta positions SAM Audio as a single unified model supporting three prompt types, usable alone or in combination:

    quillbot
  • Text prompting: Describe the sound in natural language, e.g., “dog barking” or “singing voice,” and the model separates that sound from the mixture. Text prompts are a core interaction mode, with an end-to-end example available in the open-source repo using SAMAudioProcessor and model.separate.
  • Visual prompting: Click on a person or object in a video to ask the model to isolate the audio linked to that visual object, implemented by passing video frames and masks into the processor via masked_videos.
  • Span prompting: Mark time segments where the target sound occurs; the model uses those spans to guide separation. This is crucial for ambiguous cases, such as when the same instrument appears multiple times or when a sound is brief, helping to prevent over-separation.
  • Results

    The Meta team claims SAM Audio achieves cutting-edge performance across diverse, real-world scenarios and serves as a unified alternative to single-purpose audio tools. They published a subjective evaluation across categories—General, SFX, Speech, Speaker, Music, Instr(wild), Instr(pro)—with General scores of 3.62 for sam audio small, 3.28 for sam audio base, and 3.50 for sam audio large, while Instr(pro) scores reached 4.49 for sam audio large.

    Key Takeaways

  • SAM Audio is a unified audio separation model that segments sound from complex mixtures using text prompts, visual prompts, and time span prompts.
  • The core API produces two waveforms per request: target for the isolated sound and residual for everything else, easily mapping to common edit operations like removing noise, extracting stems, or keeping ambience.
  • Meta released multiple checkpoints and variants, including sam-audio-small, sam-audio-base, sam-audio-large, plus TV variants that perform better for visual prompting. The repo also includes a subjective evaluation table by category.
  • The release includes tooling beyond inference: Meta provides a sam-audio-judge model that scores separation results against a text description, evaluating overall quality, recall, precision, and faithfulness.
  • changelly
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Fintech Fetch Editorial Team
    • Website

    Related Posts

    Following the questions where they lead | MIT News

    Following the questions where they lead | MIT News

    July 19, 2026
    Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do

    Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do

    July 18, 2026
    Examining Google DeepMind's AI bioresilience push

    Examining Google DeepMind’s AI bioresilience push

    July 17, 2026
    Thinking Machines Lab Releases Inkling: A 975B-Parameter Open-Weights Multimodal MoE With 41B Active Parameters And Controllable Thinking Effort

    Thinking Machines Lab Releases Inkling: A 975B-Parameter Open-Weights Multimodal MoE With 41B Active Parameters And Controllable Thinking Effort

    July 16, 2026
    Add A Comment

    Comments are closed.

    Join our email newsletter and get news & updates into your inbox for free.


    Privacy Policy

    Thanks! We sent confirmation message to your inbox.

    binance
    Latest Posts
    Following the questions where they lead | MIT News

    Following the questions where they lead | MIT News

    July 19, 2026
    The lazy way I make money with AI in 2026

    The lazy way I make money with AI in 2026

    July 19, 2026
    France Orders ISPs to Block Polymarket After French Traffic Surges to 578,000 Visits

    rewrite this title in other words: France Orders ISPs to Block Polymarket After French Traffic Surges to 578,000 Visits

    July 18, 2026
    Cointelegraph

    rewrite this title in other words: Bitcoin Drops Back to Its Local Range as Bear-Market History Repeats

    July 18, 2026
    Chainlink Co-Founder Nazarov Reveals 3 Trends He’s Watching Closely

    rewrite this title in other words: Chainlink Tests Support As CCIP Moves From Hype To Usage

    July 18, 2026
    quillbot
    LEGAL INFORMATION
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Top Insights
    Buy or Sell? What Michael Saylor's Cryptic New Tweet Means for Bitcoin

    rewrite this title in other words: Buy or Sell? What Michael Saylor’s Cryptic New Tweet Means for Bitcoin

    July 19, 2026
    Cointelegraph

    rewrite this title in other words: Electronic Transactions Association CEO Expecting More Partnerships with Bitcoin Startups

    July 19, 2026
    synthesia
    Facebook X (Twitter) Instagram Pinterest
    © 2026 FintechFetch.com - All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.