A group of researchers posted a preprint on arXiv on September 28 describing RAPTOR, a ridge-adaptive logistic probe that extracts concept vectors from frozen large language models.

The method aims to improve the probe-then-steer pipeline used for additive activation steering, where a learned vector is added to a layer’s representation during inference. Accuracy, directional stability under ablation, and low training cost are the stated priorities, the authors write, because vectors that degrade when ablated or cost too much to compute undermine downstream steering.

Probing — training a lightweight classifier on a model’s internal representations — has become a standard way to study what information LLMs encode. When the trained probe’s weights are normalized into a direction and injected back into the model, the operation is known as activation steering. The reliability of that pipeline rests on the quality of the vector, and existing methods often trade accuracy against stability or require substantial compute.

RAPTOR uses an L2-regularized logistic regression whose penalty strength is tuned on a validation set; the resulting weights, after normalization, serve as the concept vector. Across experiments on instruction-tuned LLMs and human-written concept datasets, the preprint reports that RAPTOR matches or exceeds strong baselines in classification accuracy while achieving competitive directional stability and substantially lower training cost. Qualitative demonstrations of downstream steering are included to support the quantitative claims.

The paper also provides a theoretical analysis using the Convex Gaussian Min-max Theorem in an idealized high-dimensional few-shot teacher-student setting, which the authors say explains how the ridge penalty mediates both probe accuracy and vector stability and aligns qualitatively with trends seen on real LLM embeddings.

The preprint has not been peer-reviewed and presents no independent test results beyond the authors’ own experiments.