
AI safety
-

Jailbreak Methods Evaluation: StrongREJECT Benchmark Insights

In the realm of AI safety, the evaluation of jailbreak methods is a critical area of investigation, particularly as advanced models like GPT-4 become increasingly prevalent.These evaluations assess how effectively certain techniques can circumvent the safeguards built into these AI systems, potentially leading to harmful prompt responses.
-

International AI Safety Report 2025

This is the first International Scientific Report on the Safety of Advanced AI. Following an interim publication in May 2024, a diverse group of 96 Artificial Intelligence (AI) experts contributed to this first full report, including an international Expert Advisory Panel nominated by 30 countries, the Organisation for Economic Co-operation and Development (OECD), the European…






