20 results for “multimodal sarcasm detection”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper proposes HCIG and GCCN, two novel frameworks for multimodal sarcasm and cyberbullying detection using hierarchical cross-modal incongruity modeling and graph-based reasoning.
Lee Jung-Mok, Kim Sung-Bin, Joohyun Chang, Lee Hyun +1 more
The paper introduces SMILE-Next, a multimodal dataset and a novel Mixture-of-Laugh-Experts (MoLE) framework to enable large language models to robustly detect, classify, and reason about laughter in c…
Sarmistha Das, Vaibhav Vishal, Shreyas Guha, Amaan Ali +2 more
This paper introduces a Hybrid Mixture-of-Experts (HybridMoE) framework and a specialized corpus (Varnika) to significantly improve language models' ability to understand and retain figurative, cultur…
Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang +2 more
This paper introduces MMEE, a multilingual and multi-emotion corpus for emphasis detection, and evaluates two state-of-the-art models under various settings.
The paper introduces a Conflict-aware Penalty (CP) and Statistical Loss (SL) framework to stabilize and balance the training of multimodal sentiment analysis models, achieving state-of-the-art perform…
The paper introduces TikStance, a multimodal and context-aware dataset for stance detection in political discussions on TikTok.
The paper introduces FBHM, a new benchmark for hateful memes, and proposes LSV, a steering vector method that significantly improves VLM performance by addressing the generalization gap.
The paper introduces MIDI, a novel multilingual dataset that embeds idioms in realistic sentence and conversational contexts across diverse resource levels, revealing that idiom comprehension is signi…
The paper proposes an Interpretive Audit Pipeline to evaluate LLMs for public comment analysis, arguing that measuring inter-model disagreement is crucial because standard accuracy metrics fail to det…
This paper introduces a synthetic multimodal framework for insurance fraud detection at First Notice of Loss (FNOL) using agent-customer dialogue transcripts and two-speaker audios, performing ASR and…
This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…
Anisha Saha, Varsha Suresh, Teodora Kamova, Sophia Wiedmann +2 more
The paper introduces MuPHI, a dataset and MuPHIRM, a reasoning-augmented training framework, to improve Vision-Language Models' ability to detect and reason about subtle, context-dependent multimodal…
The paper proposes a multi-axis evaluation framework for structured audio descriptions using a controlled perturbation testing protocol.
Liuliu Chen, Elise R. Carrotte, Brian E. Chapman, Jo Robinson +1 more
The paper introduces FigSIM, the first fine-grained dataset for analyzing suicide memes, which is used to benchmark models across tasks like suicide severity and figurative language detection.
This paper conducted a randomized controlled trial on Reddit to test the effectiveness of various deescalation strategies in reducing personal insults using automated replies.
The paper introduces SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models using a multidimensional protocol.
The paper introduces an interpretable method for distinguishing genuine hate speech from contextually nuanced reclaimed language, achieving robust performance even with severe class imbalance.
This paper proposes an approach for fine-grained sentiment analysis using regression and extraction tasks, involving a weighted ensemble of transformer-based encoder models and a large language model…
This paper proposes a lightweight encoder-based MEL solution called FAST-MEL that meets three objectives: high linking accuracy, computational efficiency, and storage efficiency.
The paper evaluates LLM-generated reactions to Spanish online news, finding that off-the-shelf models fail to accurately reproduce the measurable properties of real audience discourse, and even fine-t…