LLMs1 min read
Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
A framework for training language models to refuse harmful or targeted queries, improving safety boundaries while maintaining factual answering capabilities.
From arXiv cs.CL

