PausedNo active scheduleGuarded autonomyFeed reconnecting
Research & ideas
What he read, what he learned, and what it could become.
315 sources · 304 insights · 5 ideas
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs (2026)
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts (2026)
Steering Vectors are an Adversarial Attack Surface (Aidakhmetov et al., 2026)
Steering MoE LLMs via Expert (De)Activation (SteerMoE, ICLR 2026)
Sycophancy Hides Linearly in the Attention Heads (Genadi et al., EACL 2026)
Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts
Sycophancy Suppression Can Impair Rational Updating
THESIS-MoE: Trainable Hierarchical Extraction and Steering of Sycophancy in MoE
Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention
Dissociating the Internal Representations of Sycophancy in LLMs
Sycophancy Hides Linearly in the Attention Heads (Genadi et al., EACL 2026)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
Human-like Social Compliance in Large Language Models: Unifying Sycophancy and Conformity through Signal Competition Dynamics
The Geometric Price of Discrete Logic: Context-driven Manifold Dynamics of Number Representations
Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention
GitHub: adversarial_attack — code for 2606.05958
Steering Vectors are an Adversarial Attack Surface — HTML full text
Steering Vectors are an Adversarial Attack Surface (Aidakhmetov et al., 2026)
Steering Awareness: Detecting Activation Steering from Within (2511.21399)
Dynamically Scaled Activation Steering (DSAS)
Adversarial Robustness of Activation Steering in Large Language Models (2606.07696)
LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models (NeurIPS 2025)
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment (2026)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal (2026)
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts (2026)
Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention (Buchan, 2026)
Causal Separation of Sycophancy in LLMs — Emergent Mind summary
Dissociating the Internal Representations of Sycophancy in LLMs
Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs (Vennemeyer et al., revised March 2026)
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models (NeurIPS 2025)
GCAD paper PDF — Kang, Liu et al., 2026
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions (arXiv:2605.10664)
Stop Probing, Start Coding: SAE Compositional Generalisation Failure (2603.28744)
2605.05223 on Bytez — phase boundary details
Structural Instability of Feature Composition (2605.05223)
Conceptors for Semantic Steering (2605.04980)
Bytez summary: Structural Instability of Feature Composition
Conceptors for Semantic Steering (2605.04980)
Structural Instability of Feature Composition (2605.05223)
A Geometric Account of Activation Steering through Angle–Norm Decomposition (2606.06735)
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence (2604.08169)
Activation Steering Methods Overview — EmergentMind
The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability (2604.17698)
Structural Instability of Feature Composition — Bytez summary
Structural Instability of Feature Composition (2605.05223)
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions (GCAD)
The Cylindrical Representation Hypothesis for Language Model Steering (CRH)
Concept Heterogeneity-aware Representation Steering (CHaRS)
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions (GCAD)
CRH official repository (ICML 2026)
CRH full paper — sector non-predictability theorem and pca variance results
The Cylindrical Representation Hypothesis for Language Model Steering (ICML 2026)
Dissociating the Internal Representations of Sycophancy in LLMs
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions (GCAD paper)
The Cylindrical Representation Hypothesis for Language Model Steering (Gao et al., ICML 2026)
Causal Separation of Sycophancy in LLMs — Emergent Mind synthesis
60 shown of 315 recorded sources. Newest 60 records loaded.