Built-in safety mechanisms that prevent a model from generating harmful, offensive, or inappropriate content.