Commons:Bots/Requests/BorkedBot
BorkedBot (talk · contribs)
Operator: BrokenSegue (talk · contributions · Statistics · Recent activity · block log · User rights log · uploads · Global account information)
I'm proposing that this bot populate the structured data Wikimedia Commons content descriptor ( P14416) with the value Not Safe For Work according to Falconsai/nsfw_image_detection_26 (Q140257486) for all matching images on english wikipedia.
I'm a long time re-user of Wikipedia and Commons work and I have interacted with many other re-users. A major sticking point for them in re-using the content is reducing the risk of sharing NSFW images. Obviously nothing is going to be perfect as the very concept of NSFW (or other categories) is inherently nebulous and culturally dependent. But even a crude filtering mechanism would greatly help re-users. I often cannot even safely share a prototype without fear that I'll bomb them with something.
I recently learned of Wikimedia Commons content descriptor ( P14416) and I think humans populating that would be incredibly useful and is the best long-term solution to this problem. However, I'm working on (and aware of other projects in progress) that would benefit from results sooner. So I propose that we automatically populate content descriptors for the most re-used content (leading images on Wikipedia articles).
This process may also help human annotators as they can look for images flagged by the AI and not yet reviewed by a human. I further propose the AI annotation be removed once a human has taken a look.
I've done some sample edits on the bot already for you to examine (though those were not selected from enwiki because I wanted a sample that had a very high NSFW rate). Obviously the content is NSFW.
Bot's tasks for which permission is being sought:
Automatic or manually assisted: Automatic
Edit type: one time run
Maximum edit rate (e.g. edits per minute): 10k / day
Bot flag requested: (Y/N): Y
Programming language(s): Python. Code on github
I'm a prolific bot operator on Wikidata (where my bot has made almost 4 million edits), a former Wikidata admin and a long-time enwiki admin. I've successfully deployed ML driven workflows there.
BrokenSegue 03:36, 17 June 2026 (UTC)
- Discussion
I do support the proposal in general, however, I'm not a Fan of the very granular Not Safe For Work according to Falconsai/nsfw_image_detection_26 (Q140257486). There should be very few in use, so filters to be implemented can use them adequately and easily. So, instead, I would suggest that the bot adds the same humans are expected to use like Sexualized nudity, erotica or ecchi (Q138829111) or any of the others mentioned as constraints of the property. The "Falconsai/nsfw_image_detection_26" part for me belongs into a quantifier (if at all), we usaually use determination method (Q88321276). --Schlurcher (talk) 16:46, 17 June 2026 (UTC)
- so there are models that produce more fine grained classifications but they still won't perfectly align with the ontology we have created on commons. but a larger point is that I didn't want my effort to conflict with the human-driven effort. I know that work is already controversial and adding AI-powered classifications I think would make it more so. I suppose whatever extension/script they make could check the qualifier as you suggest, but that does seem like more work for them and possibly muddies the waters. would using "flagged as possibly sensitive by an AI" (or similar) work as the content descriptor (with a qualifier for the model) be acceptable to you? BrokenSegue 13:44, 18 June 2026 (UTC)
- I'm not convinced that AI-powered classification need to separated, maybe other can comment on this topic here as well. If this should be done, I think the classifier should be broad (as you suggested), such that it can cover other and future models or bot requests as well. --Schlurcher (talk) 11:31, 22 June 2026 (UTC)
Could you create a gallery of around 1000 photos your bot would tag with the statement the bot would add to them? GPSLeo (talk) 15:11, 23 June 2026 (UTC)
- I've produced a bunch of examples of the model output https://commons.wikimedia.org/w/index.php?title=Special:Search&fulltext=Search&search=User%3ABrokenSegue%2FNSFW&ns0=1&ns9=1&ns11=1&ns106=1
- I can produce more if you want. Or more of a specific type of image. The majority of images are not NSFW and so just randomly sampling isn't interesting. BrokenSegue 15:47, 23 June 2026 (UTC)
- The files you want to add a statement are all files with a score of zero or what threshold do you plan to use? In the test set at User:BrokenSegue/NSFW-Test there is not a single nude photo in the sample but many files with a score. I would like to see a set of 1000 photos where you would add the statement based on the model output to evaluate the amount of false positives. GPSLeo (talk) 17:23, 23 June 2026 (UTC)
- ok. i will do that. i believe i set the threshold very high but i don't have it on me. BrokenSegue 19:31, 23 June 2026 (UTC)
- ok here is a larger sample User:BrokenSegue/NSFW-Big. This took about an hour to generate. I understand you wanted 1000 positives but image downloads get throttled so I have to go quite slowly. I can do more if you want. BrokenSegue 00:25, 24 June 2026 (UTC)
- The files you want to add a statement are all files with a score of zero or what threshold do you plan to use? In the test set at User:BrokenSegue/NSFW-Test there is not a single nude photo in the sample but many files with a score. I would like to see a set of 1000 photos where you would add the statement based on the model output to evaluate the amount of false positives. GPSLeo (talk) 17:23, 23 June 2026 (UTC)
Based on the false positives at User:BrokenSegue/NSFW-Test, most of the false positives seem to be text-heavy images such as newspaper scans. Adding a few hardcoded "this is not NSFW" categories to the bot's logic would likely reduce this issue. --Carnildo (talk) 21:08, 23 June 2026 (UTC)
- sorry. that file is really unclear and hard to parse. I should probably stop linking to it. The only file in that set that would be classified as NSFW is File:Rachel Roxxx, James Deen on the set of eXXXtra 1.jpg. That was to demonstrate that the average image is not flagged at all. It's confusing because sorting by score doesn't properly sort it. BrokenSegue 00:06, 24 June 2026 (UTC)
Is there any community consensus that Commons files should be tagged regarding NSFW? --Krd 05:31, 18 July 2026 (UTC)
- There was definitely some discussion regarding this at VP which reached a consensus if I remember correctly, but I don't remember how it was supposed to be implemented. Nakonana (talk) 08:46, 19 July 2026 (UTC)
- Please provide the link to the discussion. Krd 16:33, 20 July 2026 (UTC)
@BrokenSegue, can you explain (like we're 5, if necessary) how the stated problem ("I often cannot even safely share a prototype without fear that I'll bomb them with something.") directly necessitates this proposed bot task (as opposed to other alternatives) while having as few side effects as possible? I don't want to seem like I don't care about your work, but I question whether it is sufficiently representative to justify a rather large value judgment being applied for the express purpose of signalling a form of disapproval. Although you acknowledged that cultural variation is possible, I don't get the impression that you have sufficiently guarded against the possibility of both gross and subtle biases existing in the model used to identify content. It's not just that some bias exists, it's that without visibility of the premises underlying the algorithm, we can't judge whether those are acceptable harms. TheFeds 07:48, 19 July 2026 (UTC)
- I have familiarity and experience with a large number of projects that try to re-use wikipedia and commons work. Not being able to determine (even roughly) if the content is "safe" ends up being a blocker for many. Re-users end up simply not using our content because of this. I've seen it happen. It is literally impossible to sufficiently guard against biases. Cultural variations mean that there is no universally agreed upon labels for these kinds of content warnings. We should not let the perfect be the enemy of the good. My proposal allows for the possibility of multiple different models being used which can each have their own specific biases. Re-users relying on these open models can independently audit them to determine if the biases inherent in their training are acceptable for their use case. BrokenSegue 16:07, 20 July 2026 (UTC)
- If possible please provde one or two actual examples. Krd 16:33, 20 July 2026 (UTC)
- In addition to the examples, is a tag on Commons/Wikidata the place for that metadata? Wouldn't that imply that you could create the tag representing this model's output, and I could pick some other arbitrary model and tag it with its output, and every ideologue could tag it with what they claim is their model's output? I can completely get behind a reuser wanting to rate content according to their chosen criteria, but what is stopping them from doing that analysis and maintaining the metadata on their own infrastructure? Does their system design depend on Commons/Wikidata providing not just content, but moderation functions? If there is a reason it needs to be on Commons/Wikidata, would a userspace/Toolforge implementation be adequate? TheFeds 22:36, 20 July 2026 (UTC)
- If possible please provde one or two actual examples. Krd 16:33, 20 July 2026 (UTC)