It's a wryly amusing thought, but in practice, it's not going to matter if writings about unaligned AI are present in the data set or not. In practice, instrumental convergence means that bad outcomes overlap with the natural subgoals of any sufficiently advanced system. https://en.wikipedia.org/wiki/Instrumental_convergence
AI doesn't have to have a goal of causing a problem. AI doesn't have to have had a human enabler (beyond having been built). Nothing in particular is required in order to fail catastrophically; it's the default if you don't thread the very small needle.
It's a wryly amusing thought, but in practice, it's not going to matter if writings about unaligned AI are present in the data set or not. In practice, instrumental convergence means that bad outcomes overlap with the natural subgoals of any sufficiently advanced system. https://en.wikipedia.org/wiki/Instrumental_convergence
AI doesn't have to have a goal of causing a problem. AI doesn't have to have had a human enabler (beyond having been built). Nothing in particular is required in order to fail catastrophically; it's the default if you don't thread the very small needle.
It could be. who knows
AI doesn't need to want to take over. Give it autonomy; reading that it will is probably enough.