Meta ceo Mark Zuckerberg himself knew that the books used to train the company’s AI tool were pirated, according to newly-public documents in one of the California class-action lawsuits against the tech company. Attorneys for a group of author plaintiffs (including Richard Kadrey, Sarah Silverman, Andrew Sean Greer, Ta-Nehisi Coates, and Jacqueline Woodson) assert that Zuckerberg approved the use of a dataset from shadow library LibGen to train their Llama LLM, despite concerns from other employees about it provenance. “Meta has treated the so-called ‘public availability’ of shadow datasets as a get out of jail free card, notwithstanding that internal […]
AI
Fable App Generated Bigoted Reader Summaries
Social book tracking app Fable came under fire for offensive year-end reader summaries, which were generated by AI. One such summary reads, “Your journey dives deep into the heart of Black narratives and transformative tales, leaving mainstream stories gasping for air. Don’t forget to surface for the occasional white author, okay?” Others were flagged by users for being ableist, insulting their reading choices, and recommending books from “a straight, cis white man’s perspective”. Fable replied to users online with apologies for the content, noting that “Our reader summaries are intended to both capture your reading history and playfully roast your taste, […]
UK Seeks Views on Controversial AI Opt-Out Change to Copyright Law
The UK government has begun a consultation period for a proposal that would make exceptions to copyright law, allowing tech companies to use copyrighted material to train AI unless rights holders specifically opt out. In a statement, the government said, “At present, the application of UK copyright law to the training of AI models is disputed. Rights holders are finding it difficult to control the use of their works in training AI models and seek to be remunerated for its use. AI developers are similarly finding it difficult to navigate copyright law in the UK, and this legal uncertainty is […]
Amazon Aspires (and Hires) to “Enhanc[e] the Book Publishing and Reading Experience Using Cutting-Edge AI Technology”
Amazon has a job listing for a position in India for a senior applied scientist, book content experience, as part of a “Publishing & Reading Science team,” looking for people “who are passionate about Reading and are willing to take Reading to the next level.” The goal is “leveraging advances in AI to improve the Reading experience for Kindle customers, and the Publishing experience for book content creators and distributors” by “making publishers lives easier through AI.” The corpus appears to be book content that was provided for a far different purpose (and might be restricted, depending upon your agreements): […]
Bertelsmann Expands Work with AI Voice Company ElevenLabs
AI voice and text-to-speech provider ElevenLabs issued a press release celebrating their expanding relationship with Bertelsmann and its subsidiaries. “So far, 36 Bertelsmann companies are using our technology to improve production and try new approaches. AI dubbing and multi-language audio are just the start. Together, we’re finding ways to share stories across languages and formats, reaching more people in more places.” Storytel invested in ElevenLabs and has been working with them to produce AI-narrated audiobooks since mid-2023. And Harper Collins started working with ElevenLabs earlier this year to create AI audiobooks of select foreign language backlist titles. In June, ElevenLabs […]
Microsoft Is HarperCollins’ AI Licensee
Bloomberg’s Hannah Miller and Dina Bass reported that Microsoft is the large tech company that has a licensing deal with HarperCollins for nonfiction books to train AI models, “according to a person familiar with the matter.” That person indicates, “Microsoft wants the HarperCollins books for a model that it hasn’t yet announced…. The company isn’t planning to use the content to generate new books without human authors, the person said.” Harper Collins had said previously in a statement that their licensing deal authorized “limited use of select nonfiction backlist titles for training AI models to improve model quality and performance.”