Thomas Fel (@thomasfel) Bsky

Are you at #NeurIPS2025? Check out the #KempnerInstitute’s Day 2 presentations! 💡

#AI #NeuroAI

@cpehlevan.bsky.social @kanakarajanphd.bsky.social @thomasfel.bsky.social @andykeller.bsky.social @binxuwang.bsky.social @njw.fish @yilundu.bsky.social

4 months ago 5 1 0 0

Into the Rabbit Hull – Part I - Kempner Institute This blog post offers an interpretability deep dive, examining the most important concepts emerging in one of today’s central vision foundation models, DINOv2. This blogpost is the first of a […]

🐇Into the Rabbit Hull — Part 1: A Deep Dive into DINOv2🧠
Our latest Deeper Learning blog post is an #interpretability deep dive into one of today’s leading vision foundation models: DINOv2.
📖Read now: bit.ly/4nNfq8D
Stay tuned — Part 2 coming soon.
#AI #VLMs #DINOv2

5 months ago 11 2 1 0

The Bau lab is on fire ! 😍

5 months ago 3 0 0 0

Interested in doing a PhD at the intersection of human and machine cognition? ✨ I'm recruiting students for Fall 2026! ✨

Topics of interest include pragmatics, metacognition, reasoning, & interpretability (in humans and AI).

Check out JHU's mentoring program (due 11/15) for help with your SoP 👇

5 months ago 27 15 0 1

Pleased to share new work with @sflippl.bsky.social @eberleoliver.bsky.social @thomasmcgee.bsky.social & undergrad interns at Institute for Pure and Applied Mathematics, UCLA.

Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
www.arxiv.org/pdf/2510.15987

🧵1/n

5 months ago 74 16 1 0

🧠 Thrilled to share our NeuroView with Ellie Pavlick!

"From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?"

AI foundation models are coming to neuroscience—if scaling laws hold, predictive power will be unprecedented.

But is that enough?

Thread 🧵👇

5 months ago 22 8 2 0

Thx a lot Naomi ! 🙌🥹

6 months ago 1 0 0 0

This is so cool. When you look at representational geometry, it seems intuitive that models are combining convex regions of "concepts", but I wouldn't have expected that this is PROVABLY true for attention or that there was such a rich theory for this kind of geometry.

6 months ago 33 5 2 1

That concludes this two-part descent into the Rabbit Hull.
Huge thanks to all collaborators who made this work possible — and especially to @binxuwang.bsky.social , with whom this project was built, experiment after experiment.
🎮 kempnerinstitute.github.io/dinovision/
📄 arxiv.org/pdf/2510.08638

6 months ago 5 0 0 0

If this holds, three implications:
(i) Concepts = points (or regions), not directions
(ii) Probing is bounded: toward archetypes, not vectors
(iii) Can't recover generating hulls from sum: we should look deeper than just a single-layer activations to recover the true latents

6 months ago 1 1 1 0

Synthesizing these observations, we propose a refined view, motivated by Gärdenfors' theory and attention geometry.
Activations = multiple convex hulls simultaneously: a rabbit among animals, brown among colors, fluffy among textures.

The Minkowski Representation Hypothesis.

6 months ago 2 0 1 0

Taken together, the signs of partial density, local connectedness, and coherent dictionary atoms indicate that DINO’s representations are organized beyond linear sparsity alone.

6 months ago 1 0 1 0

Can position explain this ?

We found that pos. information collapses: from high-rank to a near 2-dim sheet. Early layers encode precise location; later ones retain abstract axes.

This compression frees dimensions for features, and *position doesn't explain PCA map smoothness*

6 months ago 0 0 1 0

Patch embeddings form smooth, connected surfaces tracing objects and boundaries.

This may suggests interpolative geometry: tokens as mixtures between landmarks, shaped by clustering and spreading forces in the training objectives.

6 months ago 1 0 1 0

We found antipodal feature pairs (dᵢ ≈ − dⱼ): vertical vs horizontal lines, white vs black shirts, left vs right…

Also, co-activation statistics only moderately shape geometry: concepts that fire together aren't necessarily nearby—nor orthogonal when they don't.

6 months ago 0 0 1 0

Under the Linear Rep. Hypothesis, we'd expect Dictionary to be quasi-orthogonality.
Instead, training drives atoms from near-Grassmannian initialization to higher coherence.
Several concepts fire almost always the embedding is partly dense (!), contradicting pure sparse coding.

6 months ago 1 0 1 0

🕳️🐇Into the Rabbit Hull – Part II

Continuing our interpretation of DINOv2, the second part of our study concerns the *geometry of concepts* and the synthesis of our findings toward a new representational *phenomenology*:

the Minkowski Representation Hypothesis

6 months ago 33 9 2 1

Huge thanks to all collaborators who made this work possible, and especially to @binxuwang.bsky.social. This work grew from a year of collaboration!
Tomorrow, Part II: geometry of concepts and Minkowski Representation Hypothesis.
🕹️ kempnerinstitute.github.io/dinovision
📄 arxiv.org/pdf/2510.08638

6 months ago 0 0 0 0

Curious tokens, the registers.
DINO seems to use them to encode global invariants: we find concepts (directions) that fire exclusively (!) on registers.

Example of such concepts include motion blur detector and style (game screenshots, drawings, paintings, warped images...)

6 months ago 0 0 1 0

Now for depth estimation. How does DINO know depth?

It turns out it has discovered several human-like monocular depth cues: texture gradients resembling blurring or bokeh, shadow detectors, and projective cues.

Most units mix cues, but a few remain remarkably pure.

6 months ago 0 0 1 0

Another surprise here: the most important concepts are not object-centric at all, but boundary detectors. Remarkably, these concepts coalesce into a low-dimensional subspace within (see paper).

6 months ago 1 0 1 0

This kind of concept breaks a key assumption in interpretability: that a concept is about the tokens where it fires. Here it is the opposite—the concept is defined by where it does not fire. An open question is how models form such concepts.

6 months ago 0 0 1 0

Let's zoom in on classification.
For every class, we find two concepts: one fires on the object (e.g., "rabbit"), and another fires everywhere *except* the object -- but only when it's present!

We call them Elsewhere Concepts (credit: @davidbau.bsky.social).

6 months ago 1 0 1 0

Assuming the Linear Rep. Hypothesis, SAEs arise naturally as instruments for concept extraction, they will be our companions in this descent.
Archetypal SAE uncovered 32k concepts.

Our first observation: different tasks recruit distinct regions of this conceptual space.

6 months ago 0 0 1 0

🕳️🐇 𝙄𝙣𝙩𝙤 𝙩𝙝𝙚 𝙍𝙖𝙗𝙗𝙞𝙩 𝙃𝙪𝙡𝙡 – 𝙋𝙖𝙧𝙩 𝙄 (𝑃𝑎𝑟𝑡 𝐼𝐼 𝑡𝑜𝑚𝑜𝑟𝑟𝑜𝑤)

𝗔𝗻 𝗶𝗻𝘁𝗲𝗿𝗽𝗿𝗲𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗱𝗲𝗲𝗽 𝗱𝗶𝘃𝗲 𝗶𝗻𝘁𝗼 𝗗𝗜𝗡𝗢𝘃𝟮, one of vision’s most important foundation models.

And today is Part I, buckle up, we're exploring some of its most charming features. :)

6 months ago 36 12 2 0

Really neat, congrats !

6 months ago 1 0 0 0

Superposition disentanglement of neural representations reveals hidden alignment The superposition hypothesis states that a single neuron within a population may participate in the representation of multiple features in order for the population to represent more features than the ...

Superposition has reshaped interpretability research. In our @unireps.bsky.social paper led by @andre-longon.bsky.social we show it also matters for measuring alignment! Two systems can represent the same features yet appear misaligned if those features are mixed differently across neurons.

6 months ago 9 2 2 0

Explanations are a means to an end Modern methods for explainable machine learning are designed to describe how models map inputs to outputs--without deep consideration of how these explanations will be used in practice. This paper arg...

For XAI it’s often thought explanations help (boundedly rational) user “unlock” info in features for some decision. But no one says this, they say vaguer things like “supporting trust”. We lay out some implicit assumptions that become clearer when you take a formal view here arxiv.org/abs/2506.22740

6 months ago 30 3 2 0

Beautiful work !

6 months ago 2 0 1 0

🚨Updated: "How far can we go with ImageNet for Text-to-Image generation?"

TL;DR: train a text2image model from scratch on ImageNet only and beat SDXL.

Paper, code, data available! Reproducible science FTW!
🧵👇

📜 arxiv.org/abs/2502.21318
💻 github.com/lucasdegeorg...
💽 huggingface.co/arijitghosh/...

6 months ago 44 10 1 2

Posts by Thomas Fel