Last updated
    Common Crawl
    Non-profit

    Common Crawl

    United States flagUnited States

    "普及网络数据访问"

    Founded +2008 San Francisco, California, United States Executive Director: Rich Skrenta
    社区

    简介

    Common Crawl 是一家 501(c)(3) 非营利组织,维护着一个开放的网络爬取数据存储库。自 2008 年由计算机科学家 Gil Elbaz 创立以来,该组织系统地爬取了互联网,并将 PB 级的原始 HTML、元数据和提取的文本免费提供。数据托管在 Amazon Web Services 的公共 S3 存储桶中;用户只需支付自己的计算和存储成本。Common Crawl 的语料库是最大的公开可访问网络数据集,被研究人员、记者、初创公司和主要 AI 实验室使用。在 2020 年代,它成为包括 GPT-3、GPT-4 和 LLaMA 在内的大型语言模型的主要训练来源。2025 年 11 月,《The Markup》和《Wired》的一项调查审查了 AI 公司如何使用这些数据,引发了关于同意和版权的辩论。

    使命

    Common Crawl 的使命是通过爬取互联网并将生成的内容免费提供给任何人,来普及网络数据访问。其目标是降低网络规模数据分析的门槛,促进创新,并保存开放的网络历史记录。

    历史

    Gil Elbaz 于 2007 年开始构建第一个爬虫,他坚信开放的网络数据对健康的互联网至关重要。第一个公开爬取数据集包含约 20 亿个页面,于 2008 年发布。早期基础设施依赖于捐赠的服务器和补助金。2012 年,Common Crawl 与 Amazon Web Services 合作,将数据托管在公共 S3 存储桶中,实现了免费且可扩展的访问。News Crawl 数据集于 2015 年启动,提供了一个持续更新的新闻文章源。到 2017 年,语料库已增长到 35 亿个页面,是当时公开可用的最大网络数据集。2020 年代初的 AI 热潮极大地增加了需求;Common Crawl 成为大型语言模型的默认训练语料库。为回应 2025 年调查提出的版权和同意问题,该组织宣布计划开发一个选择退出系统和更好的来源追踪。

    知名人物

    Gil Elbaz

    Founder and Chairman of the Board · 2008–present

    Founded Common Crawl; previously co-founded Applied Semantics (AdSense) and Factual.

    Rich Skrenta

    执行董事 · 2024–present

    Former CEO of Blekko; creator of the Elk Cloner virus.

    Peter Norvig

    Former Board Member · 2010–2020

    Director of Research at Google; co-author of Artificial Intelligence: A Modern Approach.

    里程碑

    2007

    Gil Elbaz begins building the first web crawler.

    2008

    Common Crawl officially founded; first public crawl dataset released (approx. 2 billion pages).

    2012

    Partnership with Amazon Web Services; data made available via public S3 buckets.

    2015

    Launch of the News Crawl dataset.

    2017

    Dataset reaches 3.5 billion pages, becoming the largest publicly available web corpus.

    2020

    Release of CC-MAIN-2020-50 with over 50 billion URLs crawled.

    2022

    Common Crawl becomes primary training source for large language models (GPT-3, LLaMA, etc.).

    2023

    Release of CC-MAIN-2023-23 with 3.1 billion pages and 400+ terabytes of uncompressed data.

    2025

    The Markup and Wired investigation raises copyright and consent issues; Common Crawl announces plans for opt-out system.

    部门

    Crawling & Infrastructure EngineeringData Processing & QualityCommunity & OutreachAdministration & Finance

    组织信息

    成立
    +2008
    总部
    San Francisco, California, United States
    国家
    United States
    Executive Director
    Rich Skrenta
    类型
    Non-profit
    类型
    Non-profit
    成立
    +2008
    国家
    美国
    创始人
    Gil Elbaz
    Annual Budget
    $1–2 million
    员工
    5–10
    Data Size
    Petabytes

    财务

    营收
    $0.0 billion
    预算
    $1–2 million
    员工
    5–10

    官方网站

    commoncrawl.org/

    加入社区

    位粉丝正在讨论