I know very little about neural networks, other than at the conceptual level, so I could be very off here, but couldn't an algorithm be defined to look for meta signals surrounding the activation pathways which could be regularly 'audited'?
What I mean is, humans have a 6th sense that tells us when things are 'off', and require further examination. This comes from having 5 senses that are highly tuned to the world around us, and work incredibly well together. In some ways, physical hardware has limitations on it's ability to 'sense' the outside world, but in other ways, the ability to analyze and collect hard data is significantly more powerful than what humans can achieve, even at a subconscious level. Do current neural network algorithms not have this kind of failsafe?
Humans don't have a very good conceptual failsafe against known exploits. Everything from optical illusions to political messaging to closeup magic can get past our perceptual filters.
In fact optical illusions are probably the best example of this; even knowing that what you perceive is not correct doesn't change your perception.
That's a great point about human conceptual failsafes, but also, the failsafe I would be referring to in your optical illusion example would be the fact that you actually know what you perceive is incorrect, receiving that information from another data source - someone told you what you are seeing is incorrect. There is never a 100% foolproof failsafe, humans can still easily be manipulated and exploited, but manipulation doesn't work 100% across the board for humans either. Is there value in leveraging failsafes across a network? (I'm just throwing ideas out here)
What I mean is, humans have a 6th sense that tells us when things are 'off', and require further examination. This comes from having 5 senses that are highly tuned to the world around us, and work incredibly well together. In some ways, physical hardware has limitations on it's ability to 'sense' the outside world, but in other ways, the ability to analyze and collect hard data is significantly more powerful than what humans can achieve, even at a subconscious level. Do current neural network algorithms not have this kind of failsafe?