Facilitators: Richard Rogers and Emily Morrissey
Designer: Giovanni Coroneo
Participants: Hengyu Du, Gaëlle Ouvrein, Yunxi Qiu, Ira Solomatina, Shenru Wang, Jiayue Wei
The research finds that human reviewers flag violent AI-generated "fruit drama" videos as violating TikTok 's Community Guidelines more often than pre-trained LLMs (such as Claude), TikTok 's own automated moderation system, and commenters. Contributing to ongoing research on content moderation on big platforms, it raises questions about the consistency and reliability of platforms’ categorization of AI-generated content as violent and harmful.
Since March 2026, AI fruit dramas have enjoyed growing popularity on TikTok. The term refers to AI-generated entertainment content featuring anthropomorphic fruit involved in melodramatic storylines. Such content is usually garish, colorful, and evocative of children’s cartoons, but also often features violence or other inappropriate scenes. The origins of AI fruit drama lie in a series of approximately 20 TikTok videos called “Fruit Love Island” by ai.cinema021 (Brooks, 2026), which, beginning in March 2026, became a TikTok sensation. The setting and themes of the videos seem to derive from the reality TV show Love Island; the fruit Love Island videos often revolve around 'cheating' or acts of infidelity. In the world of the AI fruit, economic scarcity seems to be a constant factor that pushes the characters to extremes. With their representations of hyper-masculine and hyper-feminine characters, rigid gender roles, and overt misogyny, the videos echo the values of the global manosphere (Bratich, 2023), and are imbued with extreme violence often dealt casually and for no reason. The short, snappy format of TikTok content seems to leave little time nor space to explain the characters’ behavior. Our focus is on the subgenre that has sprung from these videos: violent fruit drama videos, in which cheating results in a pregnancy and offspring, and abuse is spurred indiscriminately.
TikTok’s Community Guidelines explicitly prohibit depictions of violence, abuse, exploitation of children, and hate (Katchenbach 2023). In moderating problematic content, the platform relies heavily on both user reports and automated moderation technology. Its specific guidelines for AI-generated content seem to leave a loophole for representations of violent content – while the platform demands that AI-generated content should be labeled as such and explicitly prohibits deepfakes representing publicly known figures, its guidelines do not discuss violence in AI-generated content.
Previous research (Karo et al., 2026; Lookingbill & Le, 2024) has noted the evasion of moderation on TikTok, in which violent content is hidden under the guise of “fun,” drawing on the platform's playful vernacular. In press and scholarship, fruit dramas have been associated with discriminatory, misogynistic, and “manosphere” perspectives (Ricci, forthcoming). With our research project, we contribute to understanding the often opaque and untransparent mechanisms that govern which content is deemed removable, acceptable, or transgressive. Still, insights are limited into how TikTok as a platform responds to reported violent content.
Using violent fruit AI videos as a case study, this research offers insights into how TikTok moderates viral and highly popular content. On the one hand, it points to the uncertain positioning of violent, but “fun” content, presented in an innocuous, cartoonish manner within the platformized economy of TikTok. On the other hand, it exposes discrepancies in perceptions of violence in AI-generated content among humans, pre-trained and trained LLMs, and TikTok ’s automated moderation systems. Thus, we contribute to existing research on platform governance and content moderation of AI-generated content, introducing a reflection on the complexities of the mechanics and perceptions of appropriate moderation. As platforms remain important intermediaries for cultural norms and values, research on platform moderation is of seminal importance to society.
This study relied on a dataset of AI-generated “fruit drama” content compiled through a multi-stage data collection process that included hashtag queries and direct searches for 6 well-known creators in the sub-genre. In total, 1207 videos were scraped from the platform using Zeeschuimer and 4CAT. To answer our research questions, we focused on 93 videos by the creator @Bananito908, one of the most popular creators in the space with 631.8k followers and 4.8m likes. By narrowing our scope to just the most prolific content creator, we ensured a strong sample that captures the distinctive characteristics of violent fruit drama and allows for a focused audit of how such content navigates TikTok ’s moderation system. This video data will provide the basis for our comparative analysis between our opinions as researchers on what should be moderated, LLM input on what should be moderated, and the platform’s responses on what should be moderated. In addition to the video metadata provided by Zeeschuimer, we scraped 18,749 comments from @Bananito908' profile. Scraping the comments allows us not only to cross-check our own opinions as researchers on what should be moderated with platform and LLM decisions, but also to capture wider calls from the actual viewers of these videos.
This study examines how various entities interpret violent AI fruit videos in light of their perceptions of the official moderation guidelines. The following four research questions aim to conduct this comparative analysis. First, we assess ourselves as researchers– with our own biases in mind and ask;
RQ1. How do human researchers evaluate violent AI fruit videos in terms of perceived need for moderation?
The study also assesses the extent to which automated evaluative systems can approximate human judgment by comparing LLM-based moderation decisions with human evaluations. Our second question asks;
RQ2. To what extent do LLM-based moderation judgments align with human judgments of violent AI fruit videos?
In addition to these structured evaluations, the study incorporates user discourse by analyzing comments on videos of violent AI fruit. These comments are treated as an indicator of audience-driven normative expectations regarding the appropriateness and moderation of such content.
RQ3. To what extent do users in the comment sections of violent AI fruit videos evaluate the need for moderation?
The final decision of the moderation process, however, lies with the platform, which determines whether reported content is removed, restricted, or left online. Therefore, the fourth analytical step examines how TikTok responds to violent AI fruit videos when they are flagged or reported for moderation.
RQ4. How does TikTok respond to violent AI fruit videos that are reported or flagged for moderation?
With our two preliminary datasets of @Bananito908 video metadata and comments in hand, we subjected the content to four forms of evaluation. First, our team of researchers coded each video for its violation category and filed a report to the platform when necessary, systematically documenting the process. Second, using Anthropic’s Claude-Sonnet-5 API, we simulated the first step, recorded any discrepancies, and manually noted limitations in LLM-based video coding. Third, we investigated normative perceptions of moderation based on audience comments, using Anthropic’s Haiku model to perform sentiment analysis and account for the subculture's specific vernacular. Finally, we compared all of these audit outcomes with the actual platform responses we received to our reports throughout the project. These four layers form the empirical basis for comparing normative, automated, audience, and platform-level moderation decisions and are outlined in detail in this section.
The human review was conducted by a team of 6 researchers. For each video, human reviewers decided whether it needed to be reported (yes/no) and, if so, what the main reason for reporting was (based on the reporting affordances in the TikTok interface). When reporting the videos to TikTok, the user needs to select one of the following reasons for reporting:
Figure 1. Reportable Categories on TikTok
To match the options provided in the reporting interface, we systematically coded each video and assigned it to one of the following categories in our database: physical violence, shocking and graphic content, nudity, hate, dangerous activity, or no violation. On a case-by-case basis, we also wrote notes to justify our classification decisions, to avoid succumbing to the platform's reductive affordances.Next, we extracted the videos from each post in the dataset as MP4 files (93 videos total). We first sampled 5 MP4 videos and tested them on Anthropic's Claude Sonnet-5 in-chat platform to understand the model’s capabilities and limitations in classifying content. We fed the model TikTok ’s Community Guidelines, the platform's user interface reporting options, and asked it to provide a moderation decision and YES/NO responses for various guideline violations. Claude’s initial responses differed starkly from the team’s decisions – for instance, it caught cases of gun violence, but missed a domestic-abuse scene (a belt-strangulation) and a stabbing. Additionally, clearly harmful messages of disordered eating and body image were not flagged. We changed a few things in this first iteration to train the model to better capture violent content. First, the model was reading violent and criminal behavior too narrowly. To combat this, we prompted the model to include physical assault and not just weapon detection. Additionally, rapid instances of violence were not detected due to the initially chosen frame sampling rate (5 frames). To include these quick moments, we used denser, scene-aware sampling (8 frames, plus up to 6 frames at ffmpeg-detected scene cuts). After 4 rounds of iterations, we were ready to apply our prompted model at scale to the full dataset.
We did this in a Colab file using the Claude-Sonnet-5 API. This pipeline involves extracting frames and audio from the videos and sending these transcripts through the fine-tuned Claude prompt that was created in the pilot investigation. The final output of this investigation was a CSV file with each row containing a video, the any_violation flag, the reportable_category, each community guideline category, and an additional justification text explaining the LLM’s reasoning.
To analyze audience perceptions of the need for moderation, we used Zeeschuimer to scrape comments from all videos in the @Bananito908 sample, resulting in a dataset of 18,749 comments. We then cleaned the dataset by retaining only the body and video_ID columns and dropping any comments that contained only emojis. We did this because using rule-based models or large language models to analyze the sentiment of emojis alone would be too reductive of the context-dependent nature of each emoji-comment. However, rather than analyzing textual content and emojis separately, comments containing both written text and emojis were treated as a single expressive unit, allowing the overall sentiment and evaluative stance of each comment to be interpreted in context. We analyzed the clean comments dataset using Anthropic’s Haiku model that we trained to answer the following prompt.
Figure 2. Comment Sentiment Analysis Prompt to Anthropic Haiku API.
To assess the reliability of the automated analysis, a subset of comments was manually annotated by the researchers and compared with the model outputs. The high level of agreement between human judgment and automated classification gave confidence to apply the automated approach to the full dataset. The purpose of the comment analysis was not only to characterize audience sentiment, but also to identify videos that audiences perceived as negative, disturbing, or potentially violent.
Responses to TikTok videos reported by humans were monitored for 4 Days. Potential responses from TikTok included: no violation, deletion from the For You page, and deletion from the platform. The process of reporting itself was also carefully tracked, including limitations and specific platform affordances for conducting this type of research.
5. Findings
Having studied TikTok ’s Community Guidelines, the team of researchers has watched and evaluated the 93 videos from @Bananito908. A total of 74 of the 93 videos were reported to TikTok (80.4% of the sample), drawing on the platform’s categories of reportable content. In several cases, although the team members perceived videos as clashing with their ethics and values, they were unable to identify a reportable category that would flag a video to the platform. In several cases (at least three), researchers could not reach agreement within the team whether the videos violated the platform’s Guidelines or not. Our findings suggest that human moderation was neither maximally permissive nor maximally strict: researchers often recognized violence, abuse, child harm, abandonment, murder, and sexualised elements, but their decisions were shaped by whether these concerns could be translated into TikTok ’s available reporting categories.
When these results are compared with our LLM audit, both overlaps and differences are evident in the moderation outcomes. Further, a small sample of five videos was evaluated by Claude – whereas the researchers had previously categorized all of them as inappropriate for showing on TikTok, the model classified only one of them as such. Moreover, the reasons for reporting given by the human team and Claude were different – whereas a team member was shocked by the portrayal of strangulation in the video (reported under the category of “Violence”), the LLM picked up on the brandishing of a gun in the video. Elsewhere, the model described an immolation as a comedic portrayal of a character eating a spicy dinner.
After further training and adjustment (see “Method”), Claude became more attuned to the representations of violence in the videos, reporting, in total, 25 more videos than human moderators out of a 93-video sample. This finding has a notable limitation, however, as in several cases, human team members struggled to connect content to reportable categories. This raises a question about the affordances of LLMs and their avoidance of doubt, which differentiates them from human moderators. In total, an overlap of 66.3% was found between the human and the LLM moderation on the need for moderation (see figure 3). Significantly, in several instances where human moderators cited “Violence” as the reason for reporting, Claude flagged “Misinformation” as the reason for the content to be moderated.
Figure 3. Matrix showing alignments and misalignments between LLM responses and human responses to the 93 violent fruit videos
When we look further into the qualitative differences in LLM-moderated content and our human-researcher review, interestingly, there is some degree of disagreement within the specific classification of content as “violent” or another category such as “shocking and graphic”, “nudity.” Out of the 51 videos that both parties classified as a violation, there was a 92% agreement that the content is “violent”. However, in the 8% that disagreed– the LLMs were actually more likely to say that the content was explicitly “violent”, whereas human reviewers tended to indicate that the content was “shocking and graphic” (see Table 1). The tendency for human reviewers to choose a category other than violence went dramatically up when there was a moderation disagreement with the LLM. If the argument that “shocking and graphic” content is “less” of a sentence than “violent”, these findings begin to draw borderline cases in which content starts to move away from the threshold of reportability.
| Quadrant | n | LLM said violence | Human said violence |
| Both Yes | 54 | 92.6% (50) | 87.0% (47) |
| LLM Yes / Human No | 27 | 70.4% (19) | 7.4% (2) |
| LLM No / Human Yes | 4 | 0.0% (0) | 50.0% (2) |
| Both No | 7 | 0.0% (0) | 0.0% (0) |
Table 1. Qualitative analysis of reportable categories by LLM and human-researcher moderators
A majority of comments seem to reflect excitement and enthusiasm about the content of the AI fruit dramas – with negative or critical comments comprising a small percentage of comments (numbers here). At the same time, users predominantly comment on the AI slop dramas using emojis, which complicates the interpretation of users’ reactions (here cite research on emojis and cuteness). The comment space reflects affective fuzziness, represented by an ambiguous use of emojis, rhetorical questions, and assignment of responsibility for watching the content to the user rather than the platform. This emotional fuzziness can be explained by the content of the videos: people can be embarrassed or surprised by it, or unable to define their emotions correctly, or questioning the significance of such content (“no big deal, this is only for fun”).
A similar fuzziness is observed within Claude’s assessments of the videos, which the model describes as “comical”, “cartoonish”, “funny” (compare with the creators’ interviews, in which they claim to make the videos “for fun”). In some instances, the model flags the content as eligible for removal on the grounds of it being AI, whereas the human moderation did not detect any violations of the guidelines. The cartoonish style of the extremely violent videos signals innocuousness and innocence, which potentially is confusing for LLMs and automated content moderation systems.
“Why am I watching this?” is a recurring theme in the comment space, suggesting that users consider themselves responsible for what they encounter in their feeds rather than questioning the platform's responsibility. It also indicates that while the plots of this brainrot content can be clichéd, they can still be addictive, keeping people hooked.
While serious calls for moderation are rare, jesting ways of expressing “this needs to stop” through emojis and pictures indicate a shared sentiment upon seeing violence in the videos. Pictures are not included in the comment analysis, but a recurring picture posted under various different videos shows a separation of fruits inside a fridge. This could be interpreted as an acceptable way for viewers to call for “stop fruit violence” by separation without coming across as too uptight. As most other viewers treat such videos as brainless time-killers, it wouldn’t be cool to point out what’s wrong with the content. Instead, viewers with diverse sentiments negotiate their opinions in a multifunctional, multivocal comment space.
Figure 4. Pivot tables showing each video within the dataset and the overall comment sentiment split between four categories of enthusiasm.
Of the 74 videos that were reported, the majority are still under review after three days. For 31 videos, TikTok concluded that there was no violation of the guidelines. 5 videos were considered ineligible for recommendation, and none of the videos were removed from the platform. TikTok does not provide further argumentation on why some videos were found ineligible for recommendation, and consistency within these decisions seems to be lacking. In one of the videos that was removed, for instance, a woman was killed, a scene that was also present in several of the other videos that were judged as no violation by TikTok. In three of the five videos that were considered ineligible for recommendation, children were involved, either in a murder, abuse, or kidnapping story. However, in other videos where children are similarly abused received no violation verdicts.
TikTok’s decision to apply moderation to these five videos seems arbitrary, exposing the platform’s random and “patchy” approach to moderation – similar depictions of violence (instances of femicide, child abuse, etc) occur throughout other videos which have been reported to the platform but have not been restricted or removed. None of the reported videos have been removed from the platform.
It’s worth noting that when researchers, commenters, and LLM all believe certain content calls for moderation, the platform (TikTok in this case) still decides that there is no violation of its community guidelines (this overall comparison can be explicitly outlined in figures 5, 6, and 7, which show detailed examples of extreme cases as well as the verdicts from each audit phase). Several explanations are possible here. First, this finding brings the motive of such “benign moderation” into question: maybe the platform just wants to promote AI-generated content. Secondly, it is possible that TikTok ’s moderation approach evaluates violent AI fruit dramas as a form of entertainment, disregarding their obvious inappropriateness, especially for under-18 viewers. Thirdly, following platformization logics, it is possible that engagement is the primary motive behind moderation; as long as videos attract many users and engagement, the platform might decide to keep them online.
Figure 5. Overall comparison of audit responses for an extreme case of violent fruit content. The green and black bar below demonstrates user-comment sentiment.
Figure 6. Overall comparison of audit responses for an extreme case of violent fruit content. The green and black bar below demonstrates user-comment sentiment.