The research investigates the impact of limiting safety constraints to particular topics within a larger subject area. It identifies that applying restrictions to specific subsets of a topic can reduce the risk of generating undesirable outputs. The research suggests that a broad restriction on a topic may unintentionally limit the model's ability to address legitimate queries or engage in relevant discussions. This targeted approach offers a potential balance between safety and functionality. The MultiverseComputingCAI project is exploring this methodology to mitigate potential harms associated with large language models.
Source: https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom