Skip to content

Community follow-up: EVAPO post-training for FireRedAudio #9

Description

@duduke321
  • Thanks for open-sourcing FireRedAudio! We built EVAPO (https://github.com/duduke321/EVAPO) to explore a question in audio QA: how can a model learn from its mistakes when answer-based RL only tells it whether a response is correct?
  • Additional audio evidence can offer richer guidance, but it may contain answer clues or mislead the teacher. It is also unavailable at inference time. EVAPO therefore first admits evidence only when, after answer clues are removed, it improves the ranking of the correct option. The teacher then evaluates the student’s own responses, adding a corrective signal only when a response has negative advantage and the teacher assigns it even lower likelihood. The student continues to answer directly from audio and the question, without privileged evidence.
  • Our best results are 83.50% on MMAU test-mini and 84.46% on MMSU.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions