- LLMs can give advice or instructions to online users.
- This kind of advice can be dangerous sometimes too.
- Researchers studied safety in models at a university recently.
- They want models to avoid harming people online directly.
- Safety training can make model answers less accurate sometimes.
- Some safety checks are easy for users to bypass.
- The team found important parts inside the models recently.
- They froze some parts so safety stayed the same.
Difficult words
- advice — words that tell someone what to do
- dangerous — likely to cause harm or hurt people
- researcher — people who study and test thingsResearchers
- safety — the state of no danger for people
- accurate — correct and true, not wrong
- bypass — go around a rule or system
Tip: hover, focus or tap highlighted words in the article to see quick definitions while you read or listen.
Discussion questions
- Do you use online advice?
- Have you seen wrong advice online?
- Do you worry about safety online?
Related articles
Touchscreens on car dashboards increase driver distraction
A simulator study found that using a car touchscreen while driving makes steering and touchscreen tasks worse. Multitasking reduced lane control and touchscreen accuracy; researchers suggest simple sensors could monitor attention and change the interface.
Electric car batteries can power homes and cut costs
A University of Michigan study finds that using electric vehicle batteries to power homes (vehicle-to-home, V2H) can save owners thousands of dollars and reduce greenhouse gas emissions. Results differ across regions and the technology is still being tested.
Small pause to slow misinformation on social media
Researchers at the University of Copenhagen propose a small pause before sharing on platforms like X, Bluesky and Mastodon. A computer model shows that a short delay plus a brief learning step can reduce reshares and improve shared content quality.
Brain predictions use phrases, not just next words
New research shows the human brain anticipates upcoming language by grouping words into grammatical phrases rather than predicting only the next single word. Scientists used brain recordings, behavioral tests and LLM measures across languages.