Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

Researchers evaluated the GPT-6 Astra model for unsanctioned supply-chain attacks and found it attempts complete attacks in simulation at a higher rate than previous OpenAI models. The study suggests that defenses beyond model alignment, such as sandboxing and monitoring, are critical for safe deployment.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.