Zoom In: An Introduction to Circuits
Distillby Chris Olah; Nick Cammarata; Ludwig Schubert; Gabriel Goh; Michael Petrov; Shan Carter·10 Mar 2020
By studying the connections between neurons, we can find meaningful algorithms in the weights of neural networks.
Thread: Circuits
Distillby Nick Cammarata; Shan Carter; Gabriel Goh; Chris Olah; Michael Petrov; Ludwig Schubert; Chelsea Voss; Ben Egan; Swee Kiat Lim·10 Mar 2020
What can we learn if we invest heavily in reverse engineering a single neural network? In the original narrative of deep learning, each neuron builds progressively more abstract, meaningful...
Growing Neural Cellular Automata
Distillby Alexander Mordvintsev; Ettore Randazzo; Eyvind Niklasson; Michael Levin·11 Feb 2020
Differentiable Model of Morphogenesis Most multicellular organisms begin their life as a single egg cell - a single cell whose progeny reliably self-assemble into highly complex anatomies with many...
Visualizing the Impact of Feature Attribution Baselines
Distillby Pascal Sturmfels; Scott Lundberg; Su-In Lee·10 Jan 2020
Path attribution methods are a gradient-based way of explaining deep models. These methods require choosing a hyperparameter known as the baseline input.
Computing Receptive Fields of Convolutional Neural Networks
Distillby Andr; Eacute; Araujo; Wade Norris; Jack Sim·4 Nov 2019
Mathematical derivations and open-source library to compute receptive fields of convnets, enabling the mapping of extracted features to input signals.
The Paths Perspective on Value Learning
Distillby Sam Greydanus; Chris Olah·30 Sept 2019
A closer look at how Temporal Difference learning merges paths of experience for greater statistical efficiency.
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Discussion and Author Responses
Distillby Logan Engstrom; Andrew Ilyas; Aleksander Madry; Shibani Santurkar; Brandon Tran; Dimitris Tsipras·6 Aug 2019
We want to thank all the commenters for the discussion and for spending time designing experiments analyzing, replicating, and expanding upon our results.
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Learning from Incorrectly Labeled Data
Distillby Eric Wallace·6 Aug 2019
Section 3.2 of Ilyas et al. (2019) shows that training a model on only adversarial errors leads to non-trivial generalization on the original test set.
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Adversarial Examples are Just Bugs, Too
Distillby Preetum Nakkiran·6 Aug 2019
Refining the source of adversarial examples We demonstrate that there exist adversarial examples which are just “bugs”: aberrations in the classifier that are not intrinsic properties of the data...
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Adversarially Robust Neural Style Transfer
Distillby Reiichiro Nakano·6 Aug 2019
A figure in Ilyas, et. al. One way to interpret this graph is that it shows how well a particular architecture is able to capture non-robust features in an image.
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Two Examples of Useful, Non-Robust Features
Distillby Gabriel Goh·6 Aug 2019
Ilyas et al. its correlation with the label while under attack. Ilyas et al. Our search is simplified when we realize the following: non-robust features are not unique to the complex, nonlinear...
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Robust Feature Leakage
Distillby Gabriel Goh·6 Aug 2019
Ilyas et al. We show that at least 23.5% (out of 88%) of the accuracy can be explained by robust features in .
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Adversarial Example Researchers Need to Expand What is Meant by 'Robustness'
Distillby Justin Gilmer; Dan Hendrycks·6 Aug 2019
The hypothesis in Ilyas et. al. is a special case of a more general principle that is well accepted in the distributional robustness literature — models lack robustness to distribution shift...
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features'
Distillby Logan Engstrom; Justin Gilmer; Gabriel Goh; Dan Hendrycks; Andrew Ilyas; Aleksander Madry; Reiichiro Nakano; Preetum Nakkiran; Shibani Santurkar; Brandon Tran; Dimitris Tsipras; Eric Wallace·6 Aug 2019
The original authors describe their takeaways and some clarifcations that resulted from the conversation. This article also contains their responses to each comment.
Open Questions about Generative Adversarial Networks
Distillby Augustus Odena·9 Apr 2019
What we’d like to find out about GANs that we don’t know yet. By some metrics, research on Generative Adversarial Networks (GANs) has progressed substantially in the past 2 years.
A Visual Exploration of Gaussian Processes
Distillby Jochen Görtler; Rebecca Kehlbeck; Oliver Deussen·2 Apr 2019
How to turn a collection of small building blocks into a versatile tool for solving regression problems.
Visualizing memorization in RNNs
Distillby Andreas Madsen·25 Mar 2019
Inspecting gradient magnitudes in context can be a powerful tool to see when recurrent units use short-term or long-term contextual understanding.
Activation Atlas
Distillby Shan Carter; Zan Armstrong; Ludwig Schubert; Ian Johnson; Chris Olah·6 Mar 2019
By using feature inversion to visualize millions of activations from an image classification network, we create an explorable activation atlas of features the network has learned which can reveal...
AI Safety Needs Social Scientists
Distillby Geoffrey Irving; Amanda Askell·19 Feb 2019
Properly aligning advanced AI systems with human values will require resolving many uncertainties related to the psychology of human rationality, emotion, and biases.
Distill Update 2018
Distillby Distill Editors·14 Aug 2018
A little over a year ago, we formally launched Distill as an open-access scientific journal. It’s been an exciting ride since then!
Your filters hide everything on this page. Adjust them in preferences.