Hacker News
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
fastball
|next
[-]
Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.
hypfer
|next
|previous
[-]
The kind where malicious intent is okay if the words are nice.
___
Or, rephrased: How big is the space in which you can tune this model without retraining.
Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _truly_ as flexible as claimed?
__
Maybe something like "Is this guy a corporate fraud that is going to waste my time with performative nonsense?"
That would be the true test for a moderation model and I would be immensely impressed if it could manage to pull that off.
___
Edit: Looking at the paper though.. probably not.
I suppose this is useful for B2B, which seems to be mistrals whole thing. Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Maybe opinions on those base datasets could occasionally differ more than the model can be steered.
pwython
|next
|previous
[-]
mosura
|next
|previous
[-]
You have to give it to Mistral they do at least know what the market near them says they want right now. The great problem is in a few years of this that market won’t be worth anything.
Edit to add, you could also add this to an AI workflow so as to produce content that walks right up to the line but doesn’t trigger it.
mcintyre1994
|root
|parent
|next
[-]
I think that's the main service that xAI provide for X.
petcat
|root
|parent
|previous
[-]
It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content.
I guess they know that the EU AI Act, Chat Control, etc are going to cause a lot of companies to need this kind of compliance.
lava_pidgeon
|root
|parent
|next
[-]
petcat
|root
|parent
[-]
Heavily moderated by humans with discretion.
Not AI chat bots following a rules engine.
LunaSea
|root
|parent
|next
[-]
AI can do exactly that.
baq
|root
|parent
|next
|previous
[-]
Pretty sure both of the above have extensive automation in their moderation.
colechristensen
|root
|parent
|previous
[-]
The correct way to moderate is automation with certainty falling back to humans with discretion.
The new frontier of moderation should be blocking illiterate comments, as in the commenter is replying as though they didn't read or read and didn't understand.
lenerdenator
|next
|previous
[-]
BlackRabbit1
|root
|parent
|next
[-]
They had kept up in the mid-range a few years ago. But this standing is sadly long gone.
If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.
yborg
|root
|parent
[-]
baq
|root
|parent
[-]
LunaSea
|root
|parent
[-]
NitpickLawyer
|root
|parent
|next
[-]
Mind sharing such cases? I'm not aware of any so far. There's the one with images, but that's commonly miss-understood, that case was ruled on a technicality (i.e. copyright needs to be attributed to a person, not a model)
elianaive
|next
|previous
[-]
petcat
|previous
[-]
"Shieldstral" is an awkward and bad name
braiamp
|root
|parent
|next
[-]
cyanregiment
|root
|parent
|next
|previous
[-]
The stral the broke the camel's back?