Generative Artificial Intelligence Development Training Data Acquisition Emergency

JK 2014 
Created at
Updated at  
7,061 0 0

Securing data, the most important element of generative artificial intelligence (AI) development, is on alert. Until now, AI developers have trained AI models by raking various data online through an automatic program (bot) that is in charge of crawling. Recently, content companies including media companies have banned access to their websites. Previously, AI companies were suing for copyright when using content without permission, but now they are blocking data collection itself.

It is important for big techs to have as much data as possible to learn to improve the performance of AI models, but there are concerns that the amount and quality of data they can use are deteriorating. The Epoch Institute, an American AI research institute, even predicted that if the current trend continues, it will be almost impossible to obtain new AI learning data between 2026 and 2032. It means that learning new data will become increasingly difficult if you do not pay for copyright properly.

Recently, the U.S. cybersecurity company Cloudflare launched a free program that prevents unauthorized data from being taken. It prohibits OpenAI, Google, and Apple from accessing the site without the consent of the website owner. The company said, "It will provide tools to prevent malicious operators from crawling websites on a large scale."

Reddit, the largest online community in the U.S. that uses a lot of AI learning, strengthened anti-crawling tools last month. Reddit signed a contract with Google to provide content to OpenAI for a fee this year, but it is trying to strictly prevent unauthorized crawling of its content.

In particular, media companies with high-quality data have already blocked data collection. According to Reuters, 638 out of 1,165 media companies, more than half of them, stopped searching sites for OpenAI, Google, and CommonCrawl, a non-profit data collection organization, as of the end of last year.

Generative Artificial Intelligence Development Training Data Acquisition Emergency

According to the Data Providence Initiative (DPI), a research institute run by MIT, 5 percent of 14,000 websites used to collect AI data blocked "crawler access" last year. In particular, 25 percent of high-quality content such as the media prohibits crawlers. The DPI said, "The number of measures to ban data collection across online websites is increasing rapidly."

With the proliferation of "crawler blocking," AI model developers are having difficulty securing data. Data is continuously emerging online, but it is not enough to keep up with the demand for data needed for AI learning. OpenAI's GPT-3 in 2020 learned about 300 billion tokens (minimum unit of sentence the AI learns). Launched three years later, GPT-4 has trained about 12 trillion tokens, which have increased 40 times. Meta's Generative AI Lama 3 has learned more than 15 trillion tokens this year. According to the Epoch Institute, GPT-5 is expected to learn about 60 trillion tokens, but even with all the high-quality data currently available, it may lack more than 10 to 20 trillion tokens for GPT-5 learning.

The New York Times said, "As media companies, creators, and copyright holders restrict data collection, AI developers who need to constantly secure high-quality data to keep AI models up to date are feeling threatened."

Generative Artificial Intelligence Development Training Data Acquisition Emergency


Before the advent of Generative AI, creators did not know how the data obtained by crawling were used. As the Generative AI craze blew, the value of data increased, and as media companies and creators also demanded fair value for it, their reluctance to crawl increased.

In recent years, criticism of AI developers' crawling behavior has been intensifying. This is because it has been revealed that AI developers have still collected data through crawling despite the growing demand for data copyright. According to a recent Business Insider, OpenAI and Antropic have been found to bypass tools that prevent crawling on websites. It has been revealed that PurpleLexity, an AI search startup invested by Amazon and Nvidia, has also collected data by bypassing the IT magazines Wired and Forbes' anti-crawling tools.


Key Takeaways: Data Scarcity and Access Issues in Generative AI Development

  • Data is becoming increasingly scarce for AI models: Content companies are actively blocking access to their websites, preventing AI developers from using their content to train models.
  • Copyright concerns are driving the shift: Media companies and creators are demanding fair compensation for their content, leading to a reluctance to allow crawling.
  • AI developers are facing a data crunch: While data is constantly being generated online, the demand for high-quality data for training powerful AI models is outpacing the supply.
  • Data acquisition is becoming more expensive: AI companies are now forced to pay for copyright access, making data acquisition more expensive and potentially limiting the scale of future models.
  • Anti-crawling tools are on the rise: Companies like Cloudflare and Reddit are implementing measures to prevent unauthorized data collection, further restricting access.
  • Ethical concerns are coming to the forefront: AI developers are facing criticism for bypassing anti-crawling tools and collecting data without proper consent.
  • The future of AI development is at stake: The scarcity of high-quality data could significantly impact the development of future AI models, particularly those requiring massive datasets for training.
  • The value of data is being recognized: The rise of generative AI has highlighted the importance of data and its economic value, leading to a shift in how data is accessed and used.
Tags AI Crawling Bots Amazonbot Anti-Crawling Tools Applebot ByteSpider CloudeBot Copyright DPI Data Providence Initiative GPTBot Generative AI GoogleOther ImageshiftBot Facebook X
Comments 0
Similar posts
  1. Rapidly Growing Humanoid Robot Industry with AI
    7,072
  2. Generative AI cannot scale without Responsible AI (RAI)
    7,134
  3. Japan's Current Status on Generative AI and Copyright: A Summary of Developments, Current Situation, and Key Issues
    7,811
  4. The Evolution and Production Reality of Agentic AI
    198
  1. Global Electronic Medicine Trends and Market Outlook 
    9,220
  2. Big Tech's AI Investments and RE100
    7,150
  3. Amazon's Return-to-Office Mandate: A Bold Move or a Step Backwards?
    10,437
  4. The Federal Reserve: The Money-Printing King
    9,637
  5. Green Premium, Eco-Certified Building
    8,757
  6. Europe is retreating from Carbon Neutrality.
    7,434
  7. Lack of Public Incinerators in Korea
    7,593
  8. The Biden administration announces $62 million in support of the growing hydrogen industry in the United States
    8,616
  9. Supply of EVs and Replacing Oil Demand
    8,098
  10. Digital Innovation Tools to Improve Health and Productivity in the Workplace
    7,162
  11. Virtual (VR), Augmented (AR) and Extended (XR) Real Market Status
    7,052
  12. Generative AI cannot scale without Responsible AI (RAI)
    7,134
  13. Harris And Trump's Position On the Future of American Science
    7,132
  14. Urban Extinction and Compact City
    10,128
  15. Rapidly Growing Humanoid Robot Industry with AI
    7,072
  16. Demand for AI and Electric-Differentiated Renewable Energy Surges
    7,183
  17. Starlink vs Cheon Beomseongjwa(千帆星座) Accelerates Competition for Low-Orbit Satellite Communications
    7,234
  18. AI and Exoskeleton Robots
    7,196
  19. Remember the book you read How does the world actually work
    8,230
  20. Extended Range Electric Vehicle (EREV) helps to reduce charging inconvenience and revolutionize long-distance driving
    8,270
  21. When is the oil peak?
    7,033
  22. Obstacles you should overcome for AdSense Approval
    7,476
  23. WiFi connection is established, but internet is not available on my Microsoft Windows Laptop - how to make it work?
    7,088
  24. My windows memory usage is 51% even though I haven't ran any app - how to optimize?
    7,400
  25. Technical Secification for Schwinn Men's Trailway 700c/28" Hybrid Bike
    7,879
  26. Microsoft's On-Device AI: Revolutionizing Smart Technology and Redefining Innovation
    7,575
  27. ChatGPT Reset command and Ignore the Previous Response feature to have a Solid Result
    7,680
  28. ChatGPT Connectors makes the results Perfect as you expected
    7,530
  29. The difference between 403 and 404 in HTTP
    8,356
  30. Elon Musk Refutes Reports on Tesla's Low-Cost EV Plans: What's Really Happening?
    7,504
Recently updated
  1. Michael Jackson's Billie Jean
    67
  2. Bootstrap vs. Tailwind CSS: Origins, Features, Pros & Cons, and How to Choose the Right Framework
    120
  3. The Complete Guide to Golang: History, Features, Real-World Uses, and Code Examples
    332
  4. Telemetry vs. Analytics: Understanding the Difference and Why It Matters
    211
  5. The Evolution and Production Reality of Agentic AI
    198
  6. How to Activate or Waive Your UIUC Student Health Insurance
    281
  7. Complete Guide to Building a Machine Learning Model
    316
  8. My life cuts at Las Vegas during Thanksgiving day holiday
    7,396
  9. The Cybercab Transformation: From Autonomous Taxi to Mobile Base Station
    342
  10. Harness vs. OpenClaw: Two Very Different "Agents"
    956
  11. Clean Python Environments: The Power of venv vs. Docker
    812
  12. What is Docker? Why is Docker also useful in a development environment?
    609
  13. UIUC 2026-2027 Academic Calendar
    1,565
  14. How to Build Llama 3 AI Apps with Python: Setup & User Prompts
    800
  15. Open-Source LLMs: The AI Revolution
    746
  16. Resume 2.0: Leveling Up for My First Software Gig
    2,190
  17. Not everyone will understand what this man just did
    1,817
  18. UIUC Dorm Guide: Find Your Perfect Fit !!
    1,579
  19. Unpacking IU's Shopper
    747
  20. Jackie Chan's Police Story: The Action Masterpiece
    646
  21. The IVE Story: Identity, 'I AM' Charts, and Influence
    947
  22. Tech Visionaries who graduated at UIUC - You are the Next Turn
    1,182
  23. Open Databases for Sex Crime Occurrences in the U.S.
    710
  24. Automatically copy text to the clipboard when dragging the mouse in the Cursor
    2,602
  25. My First Day at University of Illinois-Urvana Champaign
    1,179
  26. Sand, Sea, and a Splash of Fun at Newport Beach: A Family Adventure
    8,118
  27. Sun, Rocks, and Adventure: A Day at Joshua Tree National Park
    8,197
  28. Sipping the Stars: My Starbucks Adventure
    9,630
  29. Exciting explore at Sequoia National Park
    7,647
  30. My Life Shot at Death Valley
    1,687
  31. Ip Man fights with Muay Thai Master
    928
  32. Mad Clown - Don't Die
    1,009
  33. How to get Student Enrollment and Degree Verification at UIUC
    4,818
  34. LAX Thanksgiving Rush: A Joyful Reunion
    917
  35. ZO ZAZZ(조째즈) - Don`t you know (모르시나요) (PROD.ROCOBERRY)
    1,116
  36. FISHINGIRLS Unleashes Energetic EP 'Funiverse' Featuring Signature Track 'Fishing King'
    970
  37. 10CM - To Reach You (너에게 닿기를)
    1,151
  38. Feeling weak? Transform yourself at the UIUC ARC!
    1,567
  39. BOYNEXTDOOR - If I Say I Love You
    1,181
  40. The Future of Software Engineer - AI Engineering
    924
  41. G Dragon x Taeyang (Eyes Nose Lips, Power, Home Sweet Home, GOOD BOY) - LE GALA PIÈCES JAUNES 2025
    898
  42. Lie - Legend song by BIGBANG
    7,808
  43. Why ROLLBACK is useful when you work with Google Gemini CLI?
    826
  44. Reimbursement after Vaccination at McKinley Health Center
    1,012
  45. Gemini CLI makes a Magic! Time to speed up your app development with Google Gemini CLI!
    961
  46. Common Questions from UIUC school life in terms of CS Program
    1,113
  47. UIUC Immunization Compliance
    1,205
  48. LEE CHANHYUK's songs really resonate with my soul - Time Stop! Vivid LaLa Love, Eve, Endangered Love ...
    1,085
  49. LEE CHANHYUK - Endangered Love (멸종위기사랑)
    1,089
  50. Cupid (OT4/Twin Ver.) - LIVE IN STUDIO | FIFTY FIFTY (피프티피프티)
    863