ChatGPT’s artificial intelligence was trained on a vast amount of text from diverse sources. Training large-scale artificial intelligence models typically requires a team of researchers, engineers, and machine learning specialists collaborating to design and develop the model’s architecture, collect and prepare training data, and optimize the learning process.
OpenAI, the organization behind ChatGPT, used a combination of powerful hardware resources and advanced learning algorithms to train the GPT models. Training large-scale models can take weeks or even months, depending on the model’s complexity, the size of the dataset, and available computing resources. However, without specific data on ChatGPT’s training, I cannot provide a precise estimate of the people involved or hours spent.
Alexej Savreux, a thirty-four-year-old from Kansas City, says he has performed various types of work over the years. He has prepared sandwiches at a fast-food restaurant, worked as a custodian and scrap transporter. He has also worked in live theater technical audio.
These days, however, his work is less hands-on: he is an artificial intelligence trainer.
Savreux is part of a hidden army of freelance workers doing behind-the-scenes work teaching artificial intelligence systems how to analyze data in ways that generate types of text and images that have amazed people using recently popular products like ChatGPT. To improve AI accuracy, he has labeled photos and made predictions about what text apps should generate next.
The pay is $15 an hour or more, with no benefits. Out of the spotlight, Savreux and other contractors have spent countless hours in recent years teaching OpenAI’s systems to provide better responses in ChatGPT. Their feedback fulfills an urgent and endless need for the company and its AI competitors: to provide streams of phrases, labels, and other information that serve as training data.
“We are subordinate workers, but there would be no artificial intelligence language systems without us,” said Savreux, who has worked for tech startups including OpenAI, the San Francisco company that launched ChatGPT in November and sparked a wave of hype around generative artificial intelligence. “You can design all the neural networks you want, you can involve all the researchers you want, but without labelers, you don’t have ChatGPT. You have nothing,” Savreux said.
It is not work that will bring Savreux fame or wealth, but it is essential and often overlooked in the artificial intelligence field, where the supposed magic of a new technological frontier can obscure the work of contract workers.
“Much of the discourse around artificial intelligence is very congratulatory,” said Sonam Jindal, program lead for artificial intelligence, labor, and economy at Partnership on AI, a San Francisco-based nonprofit organization that promotes artificial intelligence research and education. “But we are missing a big part of the story: that this field still depends largely on a vast human workforce,” she said.
OpenAI, the company behind the ChatGPT chatbot, has ramped up hiring worldwide, recruiting approximately 1,000 remote contractors in the last six months in regions such as Latin America and Eastern Europe, according to sources informed on the matter. Approximately 60% of contractors were hired to do what is called “data labeling”: create large sets of images, audio recordings, and other information that can then be used to train artificial intelligence tools or autonomous vehicles.
The remaining 40% are computer programmers creating data for OpenAI’s models to learn software engineering tasks. OpenAI’s existing product, called Codex and launched in August 2021, is designed to translate natural language into code.
“A well-established company, determined to provide world-class AI technologies to make the world a better and more efficient place, is looking for a Python developer,” reads a job posting from OpenAI in Spanish, posted by an outsourcing agency.
Previously, OpenAI trained its models on code taken from GitHub, a repository site owned by its largest investor, Microsoft, which last week confirmed billions of dollars in new funding first reported by Semafor. But in this case, OpenAI appears to be building a dataset that includes not just lines of code, but also the human explanations behind them written in natural language.