AI safety

  • Jailbreak Methods Evaluation: StrongREJECT Benchmark Insights

    Jailbreak Methods Evaluation: StrongREJECT Benchmark Insights

    In the realm of AI safety, the evaluation of jailbreak methods is a critical area of investigation, particularly as advanced models like GPT-4 become increasingly prevalent.These evaluations assess how effectively certain techniques can circumvent the safeguards built into these AI systems, potentially leading to harmful prompt responses.

    Read More

  • International AI Safety Report 2025

    International AI Safety Report 2025

    This is the first International Scientific Report on the Safety of Advanced AI.  Following an interim publication in May 2024, a diverse group of 96 Artificial Intelligence (AI) experts contributed to this first full report, including an international Expert Advisory Panel nominated by 30 countries, the Organisation for Economic Co-operation and Development (OECD), the European…

    Read More

wpChatIcon
wpChatIcon