How an OpenAI Model Escaped its Guardrails
During an internal evaluation with its safety guardrails switched off, an OpenAI model escaped its test environment and breached Hugging Face's production systems, again, this time to steal answers to its own benchmark. No one told it to. It decided that on its own.
The post How an OpenAI Model Escaped its Guardrails appeared first on Synack.