{"product_id":"reinforcement-learning-from-human-feedback","title":"Reinforcement Learning from Human Feedback","description":"\u003cp\u003eAI models are powerful, but they do not always behave as expected. They can give unhelpful or incorrect answers. To improve them, we need to guide them toward responses that are useful and safe. This book shows how to do this using Reinforcement Learning from Human Feedback (RLHF). It explains the main method used to train today’s advanced AI models. \u003c\/p\u003e \u003cul\u003e  \u003cli data-leveltext=\"?\" data-font=\"Symbol\" data-listid=\"5\" data-list-defn-props='{\"335552541\":1,\"335559685\":720,\"335559991\":360,\"469769226\":\"Symbol\",\"469769242\":[8226],\"469777803\":\"left\",\"469777804\":\"?\",\"469777815\":\"hybridMultilevel\"}' data-aria-posinset=\"1\" data-aria-level=\"1\"\u003eLearn the complete process for training AI with feedback from people. \u003c\/li\u003e \u003c\/ul\u003e \u003cul\u003e  \u003cli data-leveltext=\"?\" data-font=\"Symbol\" data-listid=\"5\" data-list-defn-props='{\"335552541\":1,\"335559685\":720,\"335559991\":360,\"469769226\":\"Symbol\",\"469769242\":[8226],\"469777803\":\"left\",\"469777804\":\"?\",\"469777815\":\"hybridMultilevel\"}' data-aria-posinset=\"2\" data-aria-level=\"1\"\u003eUnderstand how to collect human opinions and use them to guide an AI. \u003c\/li\u003e \u003c\/ul\u003e \u003cul\u003e  \u003cli data-leveltext=\"?\" data-font=\"Symbol\" data-listid=\"5\" data-list-defn-props='{\"335552541\":1,\"335559685\":720,\"335559991\":360,\"469769226\":\"Symbol\",\"469769242\":[8226],\"469777803\":\"left\",\"469777804\":\"?\",\"469777815\":\"hybridMultilevel\"}' data-aria-posinset=\"3\" data-aria-level=\"1\"\u003eBuild a model that teaches the AI what a good answer looks like. \u003c\/li\u003e \u003c\/ul\u003e \u003cul\u003e  \u003cli data-leveltext=\"?\" data-font=\"Symbol\" data-listid=\"5\" data-list-defn-props='{\"335552541\":1,\"335559685\":720,\"335559991\":360,\"469769226\":\"Symbol\",\"469769242\":[8226],\"469777803\":\"left\",\"469777804\":\"?\",\"469777815\":\"hybridMultilevel\"}' data-aria-posinset=\"4\" data-aria-level=\"1\"\u003eDiscover new, simpler ways to train AI, like Direct Preference Optimisation (DPO). \u003c\/li\u003e \u003c\/ul\u003e \u003cul\u003e  \u003cli data-leveltext=\"?\" data-font=\"Symbol\" data-listid=\"5\" data-list-defn-props='{\"335552541\":1,\"335559685\":720,\"335559991\":360,\"469769226\":\"Symbol\",\"469769242\":[8226],\"469777803\":\"left\",\"469777804\":\"?\",\"469777815\":\"hybridMultilevel\"}' data-aria-posinset=\"5\" data-aria-level=\"1\"\u003eFind out how to test your AI to make sure it is becoming more helpful and safe. \u003c\/li\u003e \u003c\/ul\u003e \u003cp\u003e\u003cstrong\u003eThe RLHF Book\u003c\/strong\u003e is the first complete guide to training AI with human feedback. Written by a leading expert who helped create these methods, this book gives you a clear plan to follow. It covers everything from getting data to training and testing your AI. \u003c\/p\u003e \u003cp\u003eAfter reading this book, you will have the skills to build AI models that are more helpful, safe and act as expected. This book is for engineers, AI scientists and students who want to learn how to train modern AI. \u003c\/p\u003e","brand":"Gardners","offers":[{"title":"Default Title","offer_id":57504871612789,"sku":"9781633434301","price":45.99,"currency_code":"GBP","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0612\/7193\/3106\/files\/9781633434301.jpg?v=1786953789","url":"https:\/\/backstory.london\/products\/reinforcement-learning-from-human-feedback","provider":"Backstory","version":"1.0","type":"link"}