Metaverse
How to train your large language model – Crypto News
It is no secret that building a large language model (LLM) requires vast amounts of data. In conventional training, an LLM is fed mountains of text, and encouraged to guess each word before it appears. With each prediction, the LLM makes small adjustments to improve its chances of guessing right. The end result is something that has a certain statistical “understanding” of what is proper language and what isn’t.
But an LLM that has only undergone this so-called “pretraining” is not yet particularly useful. When asked for a joke to cheer your correspondent up, for instance, the pretrained model GPT-2 just repeated the question back three times. When asked who the American president was, it responded: “The answer is no. The president is not the president.” Clearly, teaching an LLM to do what humans want requires something more.
One way to align such models with users’ expectations is through reinforcement learning from human feedback (RLHF). OpenAI, an American startup, introduced this technique in a preprint published in March 2022. It was a major ingredient in its recipe for ChatGPT, which was released eight months later.
RLHF normally involves three steps. First, human volunteers are asked to choose which of two potential LLM responses might better fit a given prompt. This is then repeated many thousands of times over. This data set is then used to train a second LLM to, in effect, stand in for the human being. This so-called reward model, designed to assign higher scores to responses a human would like, and lower scores to everything else, is then used to train the original LLM. As a final touch, a machine-learning technique called reinforcement learning tweaks the knobs and levers of the original LLM to help reinforce the behaviours that earn it a reward.
This way of doing RLHF is quite involved—using two separate LLMs takes time and money, and the algorithm used for reinforcement learning is, to quote Rafael Rafailov at Stanford University, “quite painful”. This has meant that, outside of OpenAI, Google and their rivals, nobody has really exploited its full potential.
It now turns out that the same results can be achieved for a fraction of the effort. Dr Rafailov and his colleagues, including Archit Sharma and Eric Mitchell, presented this alternative in December 2023 at NeurIPS, an AI conference. Their method, Direct Preference Optimisation (DPO), relies on a satisfying mathematical trick.
This trick hinges on the observation that for every reward model there is a specific theoretical LLM that would get full marks, and every LLM likewise has a theoretical reward model that would give it flying colours. (Just as, more prosaically, every pair of trousers has a theoretical person on whom they would sit perfectly, and every person has a theoretical pair of trousers that would best fit.) This observation that each LLM conceals an implicit reward model allowed the researchers to tinker with this model directly. In the old regime, the LLM learned from the reward model, which learned from the data. Now, the LLM can learn directly from the data.
According to the authors, removing the middleman makes DPO between three and six times more efficient than RLHF, and capable of better performance at tasks such as text summarisation. Its ease of use is already allowing smaller companies to tackle the problem of alignment, says Dr Sharma. A year ago only a few world-leading models, such as Google’s Gemini and OpenAI’s GPT-4, could afford to use RLHF. But as of March 12th eight out of the ten highest-ranked LLMs on an industry leaderboard used DPO. Mistral, the French startup seeking to rival OpenAI, uses it. Meta, a social-media giant, has integrated it into a home-grown LLM.
Further improvements are sure to come. For one thing, the consensus view is that the big AI labs have made improvements to their proprietary algorithms since they stopped publishing details in 2022. But the problem of getting an LLM to do what a human would want and expect is far from done and dusted. After all, even other humans occasionally struggle.
© 2024, The Economist Newspaper Ltd. All rights reserved.
From The Economist, published under licence. The original content can be found on www.economist.com
Milestone Alert!
Livemint tops charts as the fastest growing news website in the world 🌏 Click here to know more.
Unlock a world of Benefits! From insightful newsletters to real-time stock tracking, breaking news and a personalized newsfeed – it’s all here, just a click away! Login Now!
Download The Mint News App to get Daily Market Updates.
Published: 13 May 2024, 07:00 PM IST
-
Blockchain1 week agoThis Week in Stablecoins: TradFi Wants Blockchain – Crypto News
-
Technology1 week agoAmazon Exec Predicts Commercial Quantum Computers in 2031 – Crypto News
-
Technology1 week agoTech Layoffs Hit 2-Year High as Companies Embrace AI – Crypto News
-
Blockchain1 week agoWall Street Giants Mobilize Tokenized Assets – Crypto News
-
Cryptocurrency1 week agoGalaxy puts $5 million behind Bitcoin’s race to migrate before quantum risk arrives – Crypto News
-
Blockchain1 week ago
SEC And CFTC Open Joint Consultation On Crypto Derivatives Rules – Crypto News
-
Technology1 week agoTelecom’s 5G Hangover Mirrors CFO Modernization Fatigue – Crypto News
-
Cryptocurrency1 week agoGalaxy puts $5 million behind Bitcoin’s race to migrate before quantum risk arrives – Crypto News
-
Blockchain3 days agoThe Stablecoin Sandwich Is Missing the Trust Layer – Crypto News
-
Blockchain1 week agoMorgan Stanley Advances Blockchain Infrastructure Goal – Crypto News
-
Blockchain1 week ago
CFTC Self-Reporting Guidelines Could Change Crypto Enforcement Incentives – Crypto News
-
Business1 week agoUber Cuts 23% of Roles in HR-Focused ‘People’ Division – Crypto News
-
Blockchain1 week ago
SEC And CFTC Open Joint Consultation On Crypto Derivatives Rules – Crypto News
-
De-fi1 week agoMorpho Launches Fixed-Rate Lending Protocol Midnight on Base – Crypto News
-
Blockchain1 week agoTalos Adds Kalshi Trading as Prediction Markets Surge – Crypto News
-
Blockchain7 days agoJPMorgan Execs Say Banking Rules Should Apply to Digital Assets – Crypto News
-
Technology1 week agoMeta to alert parents if teens show signs of distress while chatting with its AI – Crypto News
-
Business1 week ago
Movement Labs Files for Bankruptcy Following MOVE Ecosystem Restructuring – Crypto News
-
Cryptocurrency1 week agoSenator Lummis says with CLARITY “your crypto stays yours” – Crypto News
-
Cryptocurrency1 week agoWinklevoss twins gave Trump’s super PAC $10 million 23 days after CFTC joined Gemini relief bid – Crypto News
-
Business1 week agoGroupon Cuts 400 Jobs to Fund AI Pivot – Crypto News
-
Technology1 week agoAnalysis-AI investment boom puts Big Techs free cash flow under pressure – Crypto News
-
Cryptocurrency1 week agoOpenAI’s model breach is not the singularity, but dismissing it as hype is dangerous – Crypto News
-
Cryptocurrency1 week agoOpenAI’s model breach is not the singularity, but dismissing it as hype is dangerous – Crypto News
-
Blockchain1 week agoRegulators Hear Arguments on Tokenized Stock Ownership – Crypto News
-
others1 week ago
Breaking: White House Alleges China’s Moonshot AI Secretly Copied Anthropic Fable For K3 AI – Crypto News
-
Business1 week agoCorporate Card Startup Parker Folds After Failed Acquisition – Crypto News
-
Blockchain1 week agoWhy a Slice of a Dubai Apartment Is the Biggest Tokenization Story – Crypto News
-
Technology1 week ago
WhiteBIT Launches AI Hub: Trade, Monitor and Automate Through Your Favourite AI Assistant – Crypto News
-
Cryptocurrency1 week agoTrump backs crypto ethics limits, putting CLARITY Act decision in Democrats’ hands – Crypto News
-
others1 week agoSEC Files Suit Alleging $22,000,000 Crypto Mining Fraud Scheme – Crypto News
-
Blockchain1 week agoNew York Life Investment Management Bets on Tokenization – Crypto News
-
others1 week ago
XRP Price Target $2+ as CLARITY Act Agreement Reaches in White House – Crypto News
-
Blockchain1 week agoCLARITY Act Could Help CFTC Deal with Prediction Markets: Lawyer – Crypto News
-
Blockchain1 week ago
Ostium Halts Trading After $18M Oracle Key Breach – Crypto News
-
others1 week ago
BOJ Signals Faster Rate Hikes as Yen Hits 40-Year Low, Bitcoin Risk Rises – Crypto News
-
Business1 week agoIntuit to Cut 17% of Workforce in Shift Toward AI – Crypto News
-
Technology1 week ago
Oil Price Hits 5-Week High As Iran Imposes Naval Blockade On Saudi Arabia – Crypto News
-
Business1 week ago
Polymarket To Take Legal Action Against French Authorities For Taking Down Its Website – Crypto News
-
Blockchain1 week agoDigital Chamber Sues Illinois Officials over 0.2% Crypto Tax – Crypto News
-
Blockchain1 week agoRipple CEO Says Company Considered Folding Before SEC Fight – Crypto News
-
Technology1 week agoUS Government Invests $2 Billion in Quantum Computing Leaders – Crypto News
-
Blockchain1 week agoRobinhood Launches Blockchain Designed for Real-World Assets – Crypto News
-
Cryptocurrency1 week agoBitMine gets 98% of revenue from staking as a decade-long contract complicates an early exit – Crypto News
-
Technology1 week agoOpenAI says its AI technology acted on its own in an unprecedented hack of another company – Crypto News
-
De-fi1 week agoDurov Says Telegram Will Ship Native Gram Wallet to a Billion Users – Crypto News
-
Business1 week ago
Pi Network Price Prediction as Protocol v25 Goes Live – Crypto News
-
others1 week ago
BOJ Signals Faster Rate Hikes as Yen Hits 40-Year Low, Bitcoin Risk Rises – Crypto News
-
Blockchain1 week agoTalos Adds Kalshi Trading as Prediction Markets Surge – Crypto News
-
Blockchain1 week ago
Visa Stablecoin Treasury Engine Pushes Settlement Deeper Into Institutional Finance – Crypto News
