A Refusal Vector (or Refusal Direction) is a single 1-dimensional subspace vector within the residual stream activations of an aligned LLM that deterministically controls whether the model emits refusal behavior or compliance across user prompts.
Determines the context-augmented retrieval precision for red-teaming guardrails, activation steering, and ai safety evaluation; mastering Refusal Vector allows builders to feed clean database sources to models, minimizing hallucinations.
A Refusal Vector (or Refusal Direction) is a discovery in mechanistic safety showing that safety alignment in Large Language Models is often mediated by a single directional vector in the residual stream activations. Identified by researchers (Arditi et al., 2024), this 1D subspace acts as an ON/OFF switch for refusal behavior. When an LLM evaluates a prompt, high projection along this vector triggers refusal outputs. By subtracting this vector during inference (refusal vector ablation), researchers can bypass safety alignment without re-training weights, proving that alignment is often shallowly encoded in latent space.
A Refusal Vector is a specific direction in an LLM's internal activation space that dictates whether the model rejects or fulfills a request.
Yes, subtracting the refusal vector from model activations during inference causes aligned LLMs to fulfill harmful or restricted prompts without modifying model weights.
We currently have no direct coverage articles matching "Refusal Vector". Explore trending global AI topics below instead.
Avatarin integrated GPT-Realtime to deploy autonomous, low-latency conversational retail agents across commercial hubs.
AWS ML Blog details secure enterprise authentication patterns for autonomous AgentCore identity using private key JWT assertions.
CNCF engineers explain how liveness and readiness probe misconfigurations trigger unnecessary serverless pod wakeups.
UC Berkeley AI Research demonstrates K-Search automated kernel transpilation from NVIDIA CUDA to Apple MLX hardware primitives.