Best of

Financing Common Crawl

Mozilla recently published an excellent new report out about Common Crawl, the non-profit whose web crawls have played an important role in the development of numerous large language models (LLMs). Written by Stefan Baack and Mozilla Insights, the report is based on both public documents and new interviews with Common Crawl’s current director and crawl engineer, and goes into some detail about the history of the organization, and how its data is being used. ↳

Language Model Hacking

With the widespread success of language models on many tasks in a zero-shot setting, there has been a huge surge of interest among social scientists in wanting to use them to code or classify documents, sometimes in place of human annotators. Given both the freedom to specify prompts, and the lack of connection to domain-specific training data, a concern naturally arises as to how easily people can manipulate their designs to produce a desired conclusion. This had been on my mind recently, and so I was delighted to see that a couple of recent papers specifically take on this question, both of which conclude that there is indeed considerable latitude to produce a desired finding by manipulating the choices involved. These are extremely useful and important results, although they also open up some questions for me, which I wanted to think through here. ↳

The OpenAI Library

The New York Times recently ran a brief article about the reading room in OpenAI’s office in San Francisco. The article was heavy on images and light on text, but the overall theme was the tension between the company’s GPT models—which have been trained on vast swaths of human culture, and are therefore able to regurgitate, remix, and approximate it—versus embracing the design and aesthetic trappings of a traditional library reading room. The article mentioned a handful of books that could be found in the OpenAI library, but many more were clearly visible in the photographs that accompanied it. ↳

ChatGPT Prompt Speculations

In a recent tweet that went viral, Dylan Patel claimed to have discovered or revealed the ChatGPT prompt, using a simple hack. The tweet included a link to a text file on pastebin and a screenshot of that same text with newlines removed. More interestingly, the author suggested in a reply that anyone could replicate this finding, and a subsequent tweet included a video of ChatGPT generating text in response to the same trick. That, however, is where things get somewhat strange. ↳

Ubi Sunt

There is a duality at the heart of large language models. On the one hand, they are essentially a backwards-looking invention, a “cultural technology”, in the words of Alison Gopnik—algorithms which index and remix a large slice of human culture (though one that is typically heavily biased towards the recent past). On the other hand, they can often seem to be producing something entirely new, and can thereby leave many people with the impression of having a personality or even “sentience” (whatever that means exactly); in the most extreme cases, some people have apparently convinced themselves that such models are a step on the path towards some sort of successor species to humanity, a new regime of algorithmic children that will survive our own human catastrophes. Complicating matters here is the fact that the emergence of and widespread attention to these systems largely overlapped with the Covid-19 pandemic, a time in which we have all had additional reason to reflect on life, death, loss, and creation. ↳

Force of Law

I do not feel particularly well qualified to write about this, but it seems to me that recent events have dramatically and urgently gestured at the nature of the rule of law, and the tenuousness of reliable enforcement. Since coming into office, the new administration has been carrying out actions that many legal experts suggest are illegal. In several cases since then, judges have issued orders which demand a pause on particular actions. Most recently, the New York Times reports that a judge has now ruled that the White House has failed to comply with a court order. While less splashy than various overt actions, this kind of failure act seems to me like to very close to the red line suggested by various commentators, including conservatives. ↳