The paper argues that when language models are created using…

The paper argues that when language models are created using training data collected without adequate documentation (data sheets) about its composition, potential biases, and filtering methods, this leads to an ethical risk they call “Data Amplification,” where errors are replicated across the web.

When a stochastic parrot LLM generates a racist or sexist re…

When a stochastic parrot LLM generates a racist or sexist response because such content was present in its massive, uncurated web training data, this is best classified as Technical Bias (in the Friedman/Nissenbaum framework) because the underlying flaw is a failure of the algorithm to filter toxic content during the training process.