Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Check if all neurons lead to ROME #2

Open
lauritowal opened this issue Apr 11, 2023 · 1 comment
Open

Check if all neurons lead to ROME #2

lauritowal opened this issue Apr 11, 2023 · 1 comment

Comments

@lauritowal
Copy link

Using interpretability tools (e.g the causal tracing method from the ROME publication), we could check if we can figure out how truth is represented and in which neurons. We could even combine this approach with the previous idea and see if they produce the same results.

Link

@KayKozaronek
Copy link

I am interested in this, but I think that it would take some time to understand the ROME paper and the causal tracing method.

I will look into the paper on the weekend and decide whether I deem it plausible that I'd be able to implement it before the paper deadline.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
Projects
None yet
Development

No branches or pull requests

2 participants