The Potential of Vision-Language Models for Content Moderation of Children's Videos

Autor:	Ahmed, Syed Hammad, Hu, Shengnan, Sukthankar, Gita
Rok vydání:	2023
Předmět:	Computer Science - Computer Vision and Pattern Recognition Computer Science - Computers and Society Computer Science - Machine Learning Computer Science - Social and Information Networks
Druh dokumentu:	Working Paper
Popis:	Natural language supervision has been shown to be effective for zero-shot learning in many computer vision tasks, such as object detection and activity recognition. However, generating informative prompts can be challenging for more subtle tasks, such as video content moderation. This can be difficult, as there are many reasons why a video might be inappropriate, beyond violence and obscenity. For example, scammers may attempt to create junk content that is similar to popular educational videos but with no meaningful information. This paper evaluates the performance of several CLIP variations for content moderation of children's cartoons in both the supervised and zero-shot setting. We show that our proposed model (Vanilla CLIP with Projection Layer) outperforms previous work conducted on the Malicious or Benign (MOB) benchmark for video content moderation. This paper presents an in depth analysis of how context-specific language prompts affect content moderation performance. Our results indicate that it is important to include more context in content moderation prompts, particularly for cartoon videos as they are not well represented in the CLIP training data. Comment: 5 pages, 1 figure. Accepted at IEEE ICMLA 2023
Databáze:	arXiv
Externí odkaz:	http://arxiv.org/abs/2312.03936 Zobrazit plný text záznamu View this record from Arxiv