AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

1 NJU-LINK Team, Nanjing University
2 Kling Team, Kuaishou Technology
3 Institute of Automation, Chinese Academy of Sciences
* Equal Contribution    Corresponding Author
211300096@smail.nju.edu.cn · liujiaheng@nju.edu.cn
NJU-LINK Team

Abstract

Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speech, music, and sound effects co-evolve. Existing large multimodal models often fail at this relational step, treating audio and visual streams as loosely coupled observations, relying on automatic speech recognition, and under-specifying non-speech sounds and their links to visual events.

We present AVSCap, a framework for audio-visual captioning centered on explicit cross-modal event binding. First, we construct AVSCap-130K, a tri-modal training corpus generated by a decoupled-then-fused pipeline that anchors visual and acoustic evidence before composing grounded omni-modal captions. Second, we train AVSCap-7B with a two-stage strategy: supervised fine-tuning establishes baseline capabilities, while sample-efficient reinforcement learning uses hybrid rewards to optimize acoustic completeness and audio-visual synergy. Third, we introduce AVSCapBench, a benchmark that decomposes captions into visual, audio, and synergy events and evaluates them with fine-grained event recall.

Introduction

Omni-modal video captioning requires more than recognizing visible objects or transcribing speech. A high-quality caption should jointly describe scenes, actions, speech, sound effects, music, and the way these signals co-evolve over time. In practice, current LMMs often show an illusion of integration: they may describe visual events and audio streams separately, but fail to bind a sound to the corresponding visible action.

AVSCap addresses this gap with three connected components: a tri-modal training corpus, a 7B omni-modal captioner, and a benchmark that measures event-level recall across visual, audio, and audio-visual synergy dimensions.

Contributions

Overview

AVSCap evaluation protocol and bottlenecks
Evaluation protocol and core bottlenecks. AVSCapBench decomposes captions into visual, audio, and synergy events, revealing modality isolation, speech-centric bias, and weak cross-modal binding in current omni-modal captioners.

Key Ideas

AVSCap-130K

AVSCap-130K is a tri-modal training corpus built around decoupled-then-fused supervision. The pipeline first anchors visual-only and audio-only evidence, then fuses them into omni-modal captions where sound effects, speech, music, and visual events are explicitly bound through local temporal context.

AVSCap-130K data construction pipeline
AVSCap-130K data construction pipeline. Visual-only captions preserve audio anchor points, audio captions provide speech/SFX/music tags, and automated verification filters missing tags, altered audio meanings, and invalid local synergy.

Automated Verification

Model Training

AVSCap-7B is built on Qwen2.5-Omni-7B and trained in two stages. Supervised fine-tuning establishes dense omni-modal captioning from visual, audio, and fused captions, while GRPO further optimizes acoustic completeness and audio-visual event binding through hybrid rewards.

SFT and GRPO

Ablation on SFT scaling and GRPO optimization
Ablation on SFT data scaling and GRPO optimization. Reinforcement learning improves acoustic completeness and synergy beyond simply scaling SFT data.

AVSCapBench

1,226manually annotated video clips
60.1saverage duration
3event dimensions: visual, audio, synergy
3audio sub-types: speech, music, SFX
Comparison of AVSCapBench with existing benchmarks
Benchmark comparison. AVSCapBench is designed for long omni-modal captions with explicit audio sub-types and audio-visual synergy.
Statistics of AVSCapBench
AVSCapBench statistics. The benchmark covers diverse video domains, durations, caption lengths, and event counts.

Leaderboard

Complete AVSCapBench results from Appendix E.2. All values are recall percentages. The Overall column is the event-count-weighted benchmark score.

ModelVisualAudioSynergyTotal
SpeechMusicSFXOverall
Closed-Source Models
Gemini-3-Pro60.4379.8139.5227.7771.2948.8860.97
Gemini-3-Flash58.1479.7839.4632.3472.6548.9460.54
MiMo-v2-omni36.2465.6614.9817.2454.2828.7140.27
Open-Source Models
AVoCaDO-7B50.5970.4238.7119.2561.0729.1349.31
ASID-Captioner-7B47.4268.7330.5017.9159.0224.8445.94
ASID-Captioner-3B43.6366.9527.0617.3157.5321.3643.03
MiMo-v2.537.6469.2119.8718.3958.0130.4442.49
TimeChat-Captioner-7B37.5563.6044.4624.6358.5924.4541.31
Qwen3-Omni-30B-A3B-Instruct41.8549.089.348.6839.1716.1935.29
video-SALMONN-2-7B39.0546.7613.768.7136.5212.4332.02
UGC-VideoCaptioner-3B33.2421.3022.0011.4820.7710.4324.24
Qwen2.5-Omni-7B34.7813.924.027.2213.717.0021.53
MiniCPM-o-4.5-9B29.3316.6722.9012.2618.169.8721.47
Qwen2.5-Omni-3B30.8714.244.574.3712.415.5819.18
ARC-Hunyuan-Video-7B20.6816.493.931.9711.414.5214.49
HumanOmniV2-7B27.784.601.582.464.412.4214.10
MiniCPM-o-2.6-8B24.616.753.313.926.133.7813.66
AVSCap-7B (Ours)59.3369.4540.3630.8264.3057.7060.44

The Audio group follows the paper layout: Speech, Music, SFX, and audio Overall. AVSCap-7B (Ours) is bolded for clarity.

Evaluation Analysis

Beyond the main leaderboard, the paper studies how captioning quality changes with video duration, frame sampling rate, modality shielding, and judge selection. These analyses show that long videos remain difficult for open-source omni-modal models, while AVSCap-7B maintains stronger modality isolation and non-speech audio coverage.

Effect of video duration on AVSCapBench
Effect of video duration. Longer videos lead to lower total recall, especially for open-source models.
Effect of FPS on AVSCapBench
Effect of FPS. Increasing FPS improves recall for some models, but gains saturate and vary by model capability.

Modality Leakage and Judge Agreement

The modality shielding experiment asks models to describe only one modality and measures whether they leak information from the suppressed modality. AVSCap-7B achieves substantially lower leakage rates, suggesting stronger instruction-controllable modality separation. For automated evaluation, the paper compares multiple judges against human annotations and adopts Gemini-3.1-Pro as the default judge because it has the highest agreement across visual, audio, and synergy decisions.

ModelAudio Leakage ↓Visual Leakage ↓
Gemini-3.1-Pro38.0%48.0%
ASID-Captioner-7B49.0%29.0%
Qwen2.5-Omni-7B72.0%25.0%
Qwen2.5-Omni-3B73.0%29.0%
AVoCaDO-7B99.0%89.0%
AVSCap-7B (Ours)4.0%7.0%
JudgeVisualAudioSynergyTotal
Gemini-3.1-Pro0.970.950.910.94
DeepSeek-V4-Pro0.940.890.870.90
Qwen3.5-27B0.920.880.840.88

Error Analysis

Appendix H studies unmatched events from 200 randomly sampled cases. We categorize visual errors into missing or incorrect visual information, audio errors into incorrect acoustic descriptions or omissions, and synergy errors into missing relations, incorrect binding, or complete event omission.

Error distribution over unmatched events
Error distribution over unmatched events. Weaker open-source models often miss audio and synergy events entirely; stronger models more often fail through fine-grained incorrect descriptions or imperfect binding.

Case 1: Acoustic Omission Video #7

The baseline captures basic helicopter action but omits rotor noise, mechanical clunks, and orchestral music.

Ground TruthA helicopter scene with rotor SFX, mechanical clunks, orchestral music, and action-sound synchronization.
A close-up shot reveals a man with shoulder-length brown hair, wearing a red plaid shirt and blue jeans, standing on the landing skid of a dark blue helicopter. He holds onto the frame of the aircraft as intense sunlight creates a lens flare across the screen, illuminating the word "BUNDESPOLIZEI" printed in white capital letters on the side. The perspective then cuts to a high-angle aerial shot looking down at a circular, elevated helipad located next to a river, where the helicopter is positioned in the center of the pad with its rotors spinning rapidly, creating a loud, rhythmic chopping and whirring of helicopter blades (SFX), the man still visible on the exterior skid. The scene shifts to a close-up exterior shot of a blue helicopter cockpit, accompanied by the loud, rhythmic whirring of helicopter rotor blades (SFX), where a bright sun flare initially obscures the view through the glass. As the lighting shifts, a man with shoulder-length dark hair, wearing a maroon long-sleeved shirt, is revealed inside the pilot's seat, looking out with a focused, intense expression. He raises his left hand and firmly grabs the upper frame of the door, producing a mechanical clunking sound (SFX), his fingers curling around the metal structure as he exerts force to pull or brace the mechanism while the camera shakes slightly, accompanied by a tense, low-frequency orchestral action score. Elsewhere, a static shot shows a closed, dark grey metal door set within a rough, concrete wall, with a vertical pipe visible on the left side. Suddenly, the door is kicked open violently from the inside with a loud, sharp metallic crash and impact noise (SFX), swinging outward to the right as a man dressed in a dark suit and white shirt bursts through. He steps forward aggressively with his left leg, making a heavy footstep on concrete (SFX), raising a handgun in his right hand and aiming it directly ahead as he emerges into the corridor. The narrative returns to the action with a tight close-up of a muscular arm, clad in a red sleeve, firmly gripping a metal bar inside what appears to be a moving vehicle or structure. The scene abruptly cuts to Steve Rogers, wearing a grey t-shirt, sprinting with intense focus across an industrial rooftop.
Caption by Qwen2.5-Omni-7BMostly visual description; the relevant non-speech audio is missing.
The video opens with a man climbing into a helicopter marked "BUNDES POLIZ" on its side. The helicopter is parked on a helipad near a body of water. The man is wearing a brown jacket and seems to be preparing for a flight. The camera then cuts to a close-up of a man in a car, looking distressed and talking on the phone. The scene shifts to a man in a gray shirt and black pants running towards the helicopter. He grabs onto the helicopter's rotor and is lifted into the air. The helicopter spins rapidly, and the man struggles to hold on. The camera shows him from different angles as he hangs on, with the cityscape visible in the background. The helicopter eventually crashes, and the man falls to the ground. The video ends with a screen showing two characters in a fight scene, with the words "WATCH MORE" and "SUBSCRIBE" displayed.
Caption by AVSCap-7B (Ours)Includes rotor SFX, cockpit clicks, footsteps, orchestral score, and crash sounds.
A tense, rhythmic orchestral score with driving strings and percussion begins to play (Music) as a blue helicopter bearing the white text "BUNDESPOLIZEI" sits on a circular helipad beside a river and modern architecture. A man with shoulder-length dark hair and a maroon long-sleeved shirt approaches the aircraft, opens the door, and climbs into the pilot's seat. The high-pitched whine of the helicopter engine starting up fills the air (SFX). Inside the cockpit, the man reaches up to flip several switches on the overhead control panel, producing sharp clicking sounds (SFX), before looking out the window with a stern expression. Suddenly, a muscular man with short blonde hair, wearing a tight grey t-shirt and dark pants, bursts through a metal door and sprints across the rooftop towards the helipad, his footsteps thudding rapidly (SFX). As the helicopter lifts off the ground, its rotors chopping the air with a heavy, rhythmic thrum (SFX), the blonde man leaps from the edge of the platform and grabs onto the aircraft's black landing skid with both hands. Hanging suspended above the city skyline and the river below, the blonde man's arms strain under his body weight as the helicopter continues to ascend and bank sharply. Inside the cockpit, the dark-haired pilot grips the control stick and pushes it forward (SFX), causing the helicopter to tilt downward. Reacting quickly, the blonde man swings his body toward the helipad structure, releasing one hand from the skid to grab a yellow metal railing on the edge of the platform. Now stretched horizontally between the ascending helicopter and the stationary railing, his muscles bulge visibly as he attempts to anchor the aircraft. The pilot peers down through the window, his face tight with exertion, while the helicopter engine roars loudly, struggling against the man's immense grip (SFX). The orchestral music swells in intensity, matching the physical struggle (Music). A top-down view shows the helicopter hovering precariously close to the helipad, unable to break free from the man's hold. With a massive effort, the blonde man pulls his arm inward, dragging the helicopter back toward the platform. The helicopter loses its battle against the man's strength and crashes violently onto the helipad. The spinning rotor blades strike the concrete, shattering instantly and sending a massive cloud of dust, sparks, and metal debris flying across the platform with a deafening crunch and clatter (SFX). The music cuts out abruptly upon impact (Music). The blonde man is thrown to the ground, rolling through the thick grey smoke to avoid the wreckage. He pushes himself up onto his hands and knees, his face covered in dust, and stares intently toward the crashed helicopter as the debris settles. The scene then cuts to a static end screen featuring a digital painting of Iron Man and Captain America locked in combat, with bright sparks erupting where Iron Man's repulsor beam meets Captain America's shield. White text reading "SUBSCRIBE" appears on the left side of the screen, while "WATCH MORE" is displayed on the right. A new, dramatic, sweeping orchestral track begins to play over the static image (Music).
Error AnalysisSevere acoustic omission.
The baseline model suffers from severe acoustic omission, completely overlooking critical non-speech auditory cues such as the roaring helicopter rotor blades, mechanical clunking (SFX), and the dramatic background orchestral score (Music), while only capturing basic visual events.

Case 2: Incomplete Transcription Video #14

The baseline summarizes the dance visually but misses dialogue, interview speech, audience reactions, and several sound events.

Ground TruthDance routine with pop music, lyrics, interview speech, applause, laughter, and stage actions.
On a stage with a black curtain background and a wooden floor, a young dancer is in the midst of a handstand while upbeat pop music plays, specifically the song 'Call Me Maybe', with the lyrics 'I threw a wish in the well, don't ask me I'll never tell' clearly audible. Text overlays appear in the lower left corner reading 'MACKENZIE'S SOLO' in white and 'DANCE STYLE: ACRO' in orange. The dancer is wearing a vibrant two-piece costume featuring neon pink, green, and yellow ruffles, topped with a matching party hat. She lowers her legs to land on her feet, immediately standing upright and raising her arms in a 'V' shape while smiling at the audience. She then performs a front walkover, her legs splitting vertically in the air as she rotates over her hands. The view shifts to a medium close-up of a woman in the audience holding a blue pen, gazing intently toward the stage. The scene cuts back to the stage where the backdrop features a large banner reading "TALENT COMPETITION" in white capital letters against a red and blue background decorated with stars. Continuing the routine, the dancer arches her back in a bridge position as a female singer sings melodically, "Yeah..." (Speech). She executes an extreme flexibility move, bending her spine backward until her head rests against her feet (Speech: "is a..."). Sitting center stage, she pulls her knees to her chest, then rolls backward onto her shoulders, extending her legs straight up in a shoulder stand. From a kneeling position, she suddenly springs upward as upbeat pop dance music begins with the lyrics 'Let's go...' (Music), jumping high into the air. The camera cuts to the audience, showing a blonde woman, Melissa, wearing a red top and black blazer, smiling broadly and clapping. The view shifts to an interview where Melissa says in an explanatory tone, "Mackenzie is seven and before she" (Speech). A lower-third graphic identifies her as "MELISSA" with the subtitle "DANCE MOM". She says earnestly, "I want to really" (Speech). The scene abruptly cuts back to the stage performance where the young dancer executes a dynamic front aerial flip in mid-air as an upbeat reality TV instrumental plays in the background (Music). Back in the dimly lit auditorium, the blonde woman in the black jacket is captured in a close-up, beaming with a wide, toothy smile, with the sound of crowd cheering and applause mixed with the music (SFX). She finishes a cartwheel, landing on the dark stage floor with a thud (SFX). She raises her right arm and arches backward, executing a back walkover, landing on her feet with her right arm extended high in the air to audience applause and cheering (SFX). The scene cuts to two women in the front row with loud, distinct laughter from the audience members (SFX). Returning to the stage in front of the large red banner, the dancer performs a chin stand variation, balancing on her chest and hands with her legs split high in the air. On the glossy black floor, she transitions into a seated position, leaning back on her hands with her knees bent and feet flat on the floor, striking a final pose while smiling brightly toward the audience. A female audience member exclaims excitedly, "Oh my god!" (Speech), accompanied by loud applause and cheering from a crowd (SFX). The video concludes with a static promotional end card for the television show 'Dance Moms'.
Caption by Qwen2.5-Omni-7BMentions music and applause, but omits key spoken content and fine-grained audio.
The video opens with a young girl, Mackenzie, sitting on a stage. She is wearing a colorful outfit and a party hat. The text "MACKENZIE'S SOLO DANCE STYLE: ACRO" appears on the screen. Mackenzie then begins her dance routine, performing various acrobatic moves. She does a handstand, a cartwheel, and a backbend. The audience, including a woman in the audience, claps and cheers for her. The backdrop behind her reads "NATIONAL TALENT COMPETITION." The music plays in the background, adding to the excitement of the performance. The video ends with a shot of the audience clapping and a woman in a black dress standing on a stage with the words "WATCH MORE VIDEOS" and "THE SHOW MUST GO ON... DANCE MOMS" displayed on the screen.
Caption by AVSCap-7B (Ours)Captures lyrics, Melissa's interview, audience cheering, applause, and final reactions.
The scene opens on a dimly lit stage where a young girl stands in profile, wearing a vibrant two-piece dance outfit with pink, green, and yellow ruffles, topped with a yellow party hat featuring a pink pom-pom. On-screen text reads "MACKENZIE'S SOLO DANCE STYLE: ACRO". An upbeat pop song with a driving beat and female vocals begins to play (Music). Mackenzie raises her arms and executes a handstand, followed by a front walkover. The camera briefly cuts to a woman in the audience with dark hair, wearing a black top and holding a blue pen, watching intently. Returning to the stage, the backdrop is now visible, displaying "NATIONAL TALENT COMPETITION" in large letters with star graphics. Mackenzie performs a backbend, lowering herself to the floor, then rolls onto her stomach, sits up, and completes a backward roll before running forward with a smile. The music continues with lyrics about having a good time (Music). The view shifts to the audience, showing a woman in a red top smiling broadly next to a young girl. The scene then transitions to an interview setting with the woman, now wearing a beige jacket and a statement necklace, positioned in front of a rack of colorful costumes. On-screen text identifies her as "MELISSA DANCE MOM". She explains, "Mackenzie is seven, and before she just did it for fun, and now I think she wants to really be good." (Speech). The video cuts back to the stage as Mackenzie performs a leaping jump, landing gracefully. The camera briefly shows Melissa in the audience again, laughing joyfully. On stage, Mackenzie continues her routine, executing a one-handed cartwheel as the music's chorus repeats (Music). The camera cuts to two women in the audience, both laughing and enjoying the performance. Back on stage, Mackenzie performs a chin stand, her legs arched over her head, before rolling onto her stomach and sliding into a final seated pose. As she strikes her final pose, loud applause and cheering erupt from the audience (SFX). The camera shows Melissa clapping enthusiastically and exclaiming, "Oh, she did so good!" (Speech). The video concludes with a static end screen featuring a blue curtain backdrop and Dance Moms promotional text.
Error AnalysisIncomplete transcription.
The baseline model exhibits incomplete transcription, completely omitting the spoken dialogue and on-screen interview speech, including the mother's remarks and audience exclamations, thereby failing to capture crucial semantic knowledge from the audio modality.

Case 3: Temporal Alignment Failure Video #13

The baseline lists race events but fails to align the starting pistol, skate scraping, commentary, crash, and finish with the visuals.

Ground TruthSpeed skating with command speech, gunshot, blade SFX, commentary, crash, crowd roar, and finish call.
A high-angle view shows a short track speed skating starting line on an indoor ice rink where five skaters are lined up horizontally, wearing distinct aerodynamic suits and helmets. Initially, the athletes stand upright and relaxed, looking down at the ice markings as the official says authoritatively, "Ready." (Speech). Suddenly, a starting pistol fires with a loud gunshot (SFX), and all five skaters simultaneously drop their bodies low, bending their knees deeply and leaning their torsos forward into a crouched, ready position. The five skaters launch from their crouched starting positions as the commentator notes, "And they're away cleanly for" (Speech). They drive their legs outward and swing their arms vigorously to gain momentum with the rhythmic scraping and whooshing of skates on ice (SFX). The camera pans smoothly to the left, tracking a tight pack as they round a curve while the commentator states, "500 meters four and a half laps around the track." (Speech). Leading the pack is a skater wearing a red helmet and a black suit with red shoulders representing China. As they straighten out onto the track, a digital graphic overlay displays "4 LAPS TO GO" as the commentator excitedly says, "And they will be flying. Four laps to go now look at that speed 43.3 kilometers." (Speech). The skaters execute powerful crossover strides, creating rhythmic, sharp scraping sounds of blades cutting into the ice (SFX). A racer in the back of the pack loses his footing and crashes, sliding helplessly across the rink into the purple padded barriers as the commentator states, "And we have one racer down and out, the Kazakhstan is out." (Speech). The remaining three skaters continue their momentum down the straightaway to continuous crowd cheering (SFX). Approaching the finish line, the leader extends his right skate forward in a lunge amidst crowd cheering (SFX), while the trailing skater simultaneously lunges to close the gap. They cross the line nearly together as the commentator declaratively states, "Photo finish." (Speech). A digital timer reads 39.81 as the race concludes. The commentator declares, "But Wu Dajing, the definite winner of that heat" (Speech). The video concludes with a static promotional graphic for the Winter Olympics.
Caption by Qwen2.5-Omni-7BA visual race summary with almost no synchronized audio-visual binding.
The video opens with a group of speed skaters on an ice rink, all wearing helmets and racing suits. They are lined up, ready to start a race. The skaters begin to move, gliding swiftly across the ice. The camera captures their speed and agility as they navigate the track. The crowd in the stands is visible, cheering and watching intently. The on-screen text shows the names and rankings of the skaters, with "WU D." in the lead. The speed of the skaters is displayed, reaching up to 50.9 km/h. As the race progresses, the skaters continue to push themselves, with some making sharp turns and others maintaining a steady pace. The final lap approaches, and the skaters are neck and neck, with "WU D." still in the lead. The video ends with a close-up of the skaters crossing the finish line, and the on-screen text confirms "WU D." as the winner. The video then transitions to a screen promoting "The Winter Olympics" with the dates "Feb 8-25" and a call to action to "Subscribe Now."
Caption by AVSCap-7B (Ours)Binds the command, gunshot, blade scraping, commentary, crash, and photo finish to visual events.
The video begins on an indoor ice rink with five speed skaters lined up at the starting line in a crouched position. From left to right, the skaters wear black, white, blue, white, and red suits. A voice commands, "Ready" (Speech), followed immediately by the sharp crack of a starting pistol (SFX). The skaters launch forward, their blades scraping the ice (SFX). An announcer states, "And they're away cleanly for this heat number two of the men's 500 meters..." (Speech) while a timer graphic in the bottom right corner begins counting up. As the skaters navigate the first curve, the skater in red takes the lead, with the announcer continuing, "And they will be flying. Four laps to go now. Look at that speed, 43.3 kilometers" (Speech) as a ranking overlay appears on the bottom left. The skaters execute powerful crossover strides, creating a rhythmic scraping sound of blades cutting into the ice (SFX). The commentator excitedly notes, "Wu Dajing of China has the lead right behind him..." (Speech) amidst the continuous ambient roar of the cheering crowd (SFX). As the skaters round a curve, a racer in the back of the pack loses his footing and crashes, sliding helplessly into the padded barriers (SFX). The commentator notes, "And we have one racer down and out, the Kazakhstan is out." (Speech). The remaining three skaters maintain high speed down the straightaway to the crowd's loud cheering (SFX). Approaching the final lap, the leader maintains a narrow lead over his opponent. He extends his right skate forward in a lunge to cross the finish line (SFX) as the commentator declares, "Photo finish." (Speech). A final results graphic lists Wu Dajing first, while the commentator states, "But Wu Dajing, the definite winner of that heat" (Speech). The video concludes with a static Winter Olympics promotional screen (Music).
Error AnalysisTemporal alignment failure.
The baseline model fails to capture temporal alignment, listing isolated visual actions such as the skaters starting and crossing the finish line without syncing them with auditory triggers, such as the starting pistol gunshot (SFX) or the commentator's live voiceover (Speech).

Conclusion

AVSCap is a unified framework for fine-grained audio-visual synergy that addresses modality isolation and speech-centric bias in omni-modal video captioning. We construct AVSCap-130K, a tri-modal training corpus enforcing isolated unimodal perception before cross-modal grounding, and train AVSCap-7B with a two-stage SFT-GRPO paradigm to optimize event binding. We further introduce AVSCapBench, a human-curated benchmark with a fine-grained, event-based matching protocol for evaluating visual, audio, and synergy dimensions. Experiments show that AVSCap-7B substantially outperforms open-source baselines and approaches commercial-grade performance.

Citation

@article{wang2026avscap,
  title   = {AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning},
  author  = {Wang, Yanghai and Wang, Jiahao and Tang, Jiafu and Zhang, Yuanxing and Cao, Zhe and Bian, Hanyan and Zhang, Zijie and Luo, Weiliang and Pan, Zhiyu and Dong, Zixuan and Liu, Jiaheng and Zhang, Zhaoxiang},
  journal = {arXiv preprint arXiv:2607.12820},
  year    = {2026}
}