讀懂AI時代 Read AI Time讀懂AI時代 Read AI Time简体

訂閱《讀懂AI時代》

每早 8 點一封,昨天 AI 圈大事一次讀完。免費。

過去十幾年,AI 之所以突飛猛進,底層是三件東西在指數級變大:訓練用的算力、模型的引數量、喂進去的資料量。這也是「規模定律」和這輪 AI 資本開支狂飆的底層邏輯。 下圖收錄了 535 個知名 AI 模型——單看訓練算力,從最早的神經網路到當下的前沿大模型,已增長約 25 個數量級(1025 倍)。以下只陳述已公開的客觀資料,不預測、不構成任何投資建議。

訓練算力趨勢

訓練一個模型消耗的總計算量(浮點運算次數,10^15 FLOP = 1 petaFLOP)。縱軸為對數刻度——每上一格代表大 10 倍。 最新的前沿模型訓練算力已達約 5.0×10¹¹ petaFLOP 量級。

10⁻¹²10⁻⁸10⁻⁴110⁴10⁸10¹²1960198020002020釋出年份訓練算力(petaFLOP,對數軸)Theseus|1950|4.0×10^-14 petaFLOPPerceptron Mark I|1957|6.9×10^-10 petaFLOPPandemonium (morse)|1959|6.0×10^-7 petaFLOPSamuel Neural Checkers|1959|4.3×10^-7 petaFLOPPerceptron (1960)|1960|7.2×10^-7 petaFLOPADALINE|1960|6.6×10^-12 petaFLOPLinear Decision Functions|1962|1.6×10^-9 petaFLOPPrint Recognition Logic|1963|2.2×10^-8 petaFLOPHeuristic Reinforcement Learning|1965|1.1×10^-9 petaFLOPLTE speaker verification system|1966|1.1×10^-7 petaFLOPCognitron|1975|5.2×10^-9 petaFLOPNeocognitron|1980|2.7×10^-7 petaFLOPASE+ACE|1983|3.2×10^-7 petaFLOPDistributed representation NN|1986|3.9×10^-7 petaFLOPMLP with back-propagation|1986|6.7×10^-7 petaFLOPNetTalk (dictionary)|1987|2.8×10^-5 petaFLOPNetTalk (transcription)|1987|2.8×10^-5 petaFLOPTranslation-invariant MLP|1987|1.8×10^-5 petaFLOPMLN-ASR|1988|3.0×10^-7 petaFLOPInvariant image recognition|1989|2.7×10^-5 petaFLOPHandwritten digit recognition network|1989|1.8×10^-4 petaFLOPSpeaker-independent vowel classification|1989|7.5×10^-6 petaFLOPZip CNN|1989|1.5×10^-3 petaFLOPNETtalk reimplementation|1990|3.6×10^-5 petaFLOPBankruptcy-NN|1990|3.1×10^-6 petaFLOPSexNet compression|1990|7.9×10^-5 petaFLOPWeight Decay|1991|7.6×10^-5 petaFLOPTD-Gammon|1992|1.8×10^-2 petaFLOPCancer drug mechanism prediction|1992|5.4×10^-8 petaFLOPSiamese-TDNN|1993|1.3×10^-2 petaFLOPANN Eye Tracker|1993|1.7×10^-5 petaFLOPCeramic-MLP|1994|4.5×10^-6 petaFLOPJPMAX|1994|8.1×10^-8 petaFLOPMixture of linear models|1994|4.5×10^-4 petaFLOPNeuroChess|1994|8.6×10^-4 petaFLOPPredictive Coding NN|1994|1.9×10^-2 petaFLOPLISSOM|1995|2.0×10^-4 petaFLOPMUSIC perceptron|1996|8.8×10^-4 petaFLOPSystem 11|1996|2.6×10^-5 petaFLOPSOM-CNN|1997|3.1×10^-5 petaFLOPLSTM|1997|3.2×10^-2 petaFLOPLeNet-5|1998|2.8×10^-3 petaFLOPRECONTRA-categorized|1999|8.0×10^-3 petaFLOPRECONTRA-uncategorized|1999|3.9×10^-3 petaFLOPNeural LM|2000|6.3×10^0 petaFLOPPoE MNIST|2000|5.2×10^-2 petaFLOPDecision tree (classification)|2001|6.3×10^-2 petaFLOPNPLM (AP News)|2003|1.7×10^0 petaFLOPNPLM (Brown)|2003|1.3×10^-1 petaFLOPInvariant CNN|2004|9.7×10^-4 petaFLOPLMICA|2004|2.8×10^0 petaFLOPHierarchical LM|2005|1.2×10^-1 petaFLOPRankNet|2005|3.5×10^-3 petaFLOPSVM-CNN|2006|7.4×10^-1 petaFLOPKN-LM|2007|7.7×10^2 petaFLOPSB-LM|2007|1.5×10^3 petaFLOPGNN|2008|1.6×10^-6 petaFLOPGPU DBNs|2009|1.0×10^0 petaFLOPTwo Stage Feature Extraction (MNIST)|2009|2.1×10^-2 petaFLOPLCNP LabelMe|2009|3.3×10^0 petaFLOPLCNP MNIST|2009|4.2×10^0 petaFLOPLCNP NORB|2009|2.5×10^0 petaFLOPFeedforward NN|2010|3.5×10^-1 petaFLOPiCCCP|2010|1.1×10^0 petaFLOPPooling CNN (Caltech 101)|2010|1.2×10^0 petaFLOPPooling CNN (NORB)|2010|1.5×10^0 petaFLOPRNN LM|2010|5.4×10^1 petaFLOPDeep Autoencoders|2011|3.7×10^1 petaFLOPHigh Performance CNN (NORB)|2011|2.6×10^1 petaFLOPCNN Committee (MNIST)|2011|5.2×10^1 petaFLOPCNN Committee (NIST)|2011|2.6×10^1 petaFLOPCNN committee (traffic sign)|2011|9.9×10^-1 petaFLOPDropout (CIFAR)|2012|4.3×10^0 petaFLOPDropout (ImageNet)|2012|2.7×10^2 petaFLOPDropout (MNIST)|2012|6.0×10^0 petaFLOPUnsupervised High-level Feature Learner|2012|6.0×10^2 petaFLOPLSTM LM|2012|1.7×10^1 petaFLOPAlexNet|2012|4.7×10^2 petaFLOPDNN EM segmentation|2012|4.8×10^2 petaFLOPDistBelief Speech|2012|3.1×10^2 petaFLOPDistBelief NNLM|2013|2.6×10^3 petaFLOPReLU-Speech|2013|1.3×10^2 petaFLOPHierarchical Scene Labeling (Stanford Background)|2013|2.4×10^2 petaFLOPRCTM|2013|9.3×10^0 petaFLOPRNTN|2013|1.4×10^1 petaFLOPWord2Vec (large)|2013|3.9×10^1 petaFLOPVisualizing CNNs|2013|5.3×10^2 petaFLOPTransE|2013|1.3×10^3 petaFLOPDQN|2013|2.9×10^0 petaFLOPImage generation|2013|4.7×10^-1 petaFLOPGANs|2014|5.2×10^2 petaFLOPSPPNet|2014|3.4×10^3 petaFLOPSmooCT|2014|6.9×10^1 petaFLOPACF-WIDER|2014|7.6×10^-2 petaFLOPRNNsearch-50*|2014|1.6×10^3 petaFLOPVGG16|2014|1.2×10^4 petaFLOPVGG19|2014|1.1×10^4 petaFLOPSeq2Seq LSTM|2014|5.6×10^4 petaFLOPSPN-4+KN5|2014|4.4×10^1 petaFLOPGoogLeNet / InceptionV1|2014|1.5×10^3 petaFLOPTA-CNN|2014|1.1×10^1 petaFLOPSNM-skip|2014|3.0×10^5 petaFLOPFractional Max-Pooling|2014|1.0×10^2 petaFLOPADAM (CIFAR-10)|2014|6.3×10^-1 petaFLOPMSRA (C, PReLU)|2015|2.4×10^4 petaFLOPgenCNN + dyn eval|2015|3.4×10^1 petaFLOPTC-DNN-BLSTM-DNN|2015|1.9×10^2 petaFLOPU-Net|2015|5.1×10^1 petaFLOPDCNN|2015|4.8×10^2 petaFLOPAlphaGo Fan|2015|3.8×10^5 petaFLOPSAF R-CNN|2015|1.2×10^4 petaFLOPInception v3|2015|1.0×10^5 petaFLOPResNet-101 (ImageNet)|2015|7.0×10^3 petaFLOPResNet-152 (ImageNet)|2015|1.0×10^4 petaFLOPVariational (untied weights, MC) LSTM (Large)|2015|5.9×10^0 petaFLOPAlphaGo Lee|2016|1.9×10^6 petaFLOPNamed Entity Recognition model|2016|9.7×10^1 petaFLOPR-FCN|2016|7.2×10^2 petaFLOPResNet-200|2016|3.0×10^4 petaFLOPGNMT|2016|6.6×10^6 petaFLOPPointer Sentinel-LSTM (medium)|2016|7.5×10^0 petaFLOPXception|2016|4.4×10^5 petaFLOPSPIDER2|2016|1.8×10^1 petaFLOPBIDAF|2016|3.5×10^3 petaFLOPNAS with base 8 and shared embeddings|2016|1.1×10^1 petaFLOPNASv3 (CIFAR-10)|2016|2.2×10^6 petaFLOPVD-LSTM+REAL Large|2016|2.1×10^1 petaFLOPResNeXt-101 (64×4d)|2016|1.2×10^4 petaFLOPPolyNet|2016|6.4×10^4 petaFLOPHR-ResNet101|2016|7.1×10^3 petaFLOPEnhanceNet|2016|1.3×10^2 petaFLOPDeepStack|2017|1.5×10^4 petaFLOPMoE-Multi|2017|9.4×10^4 petaFLOPTransformer (2017)|2017|7.4×10^3 petaFLOPDeepLoc|2017|5.8×10^2 petaFLOPJFT|2017|8.4×10^5 petaFLOPConvS2S (ensemble of 8 models)|2017|5.6×10^4 petaFLOPAWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2)|2017|3.0×10^2 petaFLOPRetinaNet-R101|2017|2.1×10^3 petaFLOPOpenAI TI7 DOTA 1v1|2017|6.1×10^5 petaFLOPEI-REHN-1000D|2017|1.1×10^1 petaFLOPLibratus|2017|5.5×10^5 petaFLOPGL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2)|2017|4.6×10^2 petaFLOPPyramidNet|2017|2.3×10^0 petaFLOPISS|2017|3.4×10^0 petaFLOPAWD-LSTM+WT+Cache+IOG (WT2)|2017|3.2×10^0 petaFLOPAlphaGo Zero|2017|6.5×10^5 petaFLOPAlphaGo Master|2017|3.4×10^5 petaFLOPFraternal dropout + AWD-LSTM 3-layer (WT2)|2017|3.1×10^2 petaFLOPAWD-LSTM-MoS + dynamic evaluation (WT2, 2017)|2017|3.4×10^3 petaFLOPAlphaZero|2017|1.1×10^5 petaFLOPELMo|2018|3.3×10^0 petaFLOPQRNN|2018|6.9×10^2 petaFLOPIMPALA|2018|1.7×10^5 petaFLOP4 layer QRNN (h=2500)|2018|5.9×10^2 petaFLOPYOLOv3|2018|1.3×10^4 petaFLOPDropout-LSTM+Noise(Bernoulli) (WT2)|2018|1.3×10^2 petaFLOPResNeXt-101 32x48d|2018|8.7×10^6 petaFLOPaLSTM(depth-2)+RecurrentPolicy (WT2)|2018|7.3×10^1 petaFLOPGPT-1|2018|1.8×10^4 petaFLOPFTW (For The Win)|2018|3.5×10^4 petaFLOPBig-Little Net|2018|2.5×10^2 petaFLOPBig-Little Net (speech)|2018|4.3×10^2 petaFLOPBig Transformer for Back-Translation|2018|4.8×10^5 petaFLOP(ensemble): AWD-LSTM-DOC (fin) × 5 (WT2)|2018|6.7×10^2 petaFLOPTransformer + Simple Recurrent Unit|2018|1.1×10^4 petaFLOPLSTM+NeuralCache|2018|9.8×10^-1 petaFLOPTransformer (Adaptive Input Embeddings) WT103|2018|4.5×10^4 petaFLOPBERT-Large|2018|2.9×10^5 petaFLOPTrellisNet|2018|2.8×10^3 petaFLOPMesh-TensorFlow Transformer 2.9B (translation)|2018|6.8×10^4 petaFLOPMesh-TensorFlow Transformer 4.9B (language)|2018|1.6×10^5 petaFLOPFine-tuned-AWD-LSTM-DOC (fin)|2018|5.2×10^1 petaFLOPMulti-cell LSTM|2018|2.0×10^0 petaFLOPStyleGAN|2018|3.9×10^1 petaFLOPTransformer-XL (257M)|2019|3.8×10^5 petaFLOPHanabi 4 player|2019|4.3×10^3 petaFLOPGPT-2 (1.5B)|2019|1.9×10^6 petaFLOPKataGo|2019|2.3×10^4 petaFLOPSciBERT|2019|8.9×10^4 petaFLOPCross-lingual alignment|2019|2.6×10^3 petaFLOPWeNet (Penn Treebank)|2019|7.3×10^2 petaFLOPBERT-Large-CAS (PTB+WT2+WT103)|2019|1.5×10^5 petaFLOPMuseNet|2019|2.2×10^5 petaFLOPAWD-LSTM-DRILL + dynamic evaluation† (WT2)|2019|4.1×10^2 petaFLOPDLRM-2020|2019|4.0×10^3 petaFLOPXLNet|2019|6.2×10^6 petaFLOPTransformer-XL Large + Phrase Induction|2019|3.8×10^5 petaFLOPAWD-LSTM + MoS + Partial Shuffled|2019|3.2×10^2 petaFLOPRoBERTa Large|2019|8.5×10^6 petaFLOPPluribus|2019|6.6×10^1 petaFLOPtrRosetta|2019|3.8×10^4 petaFLOPUDSMProt|2019|6.4×10^2 petaFLOPMegatron-BERT|2019|2.2×10^7 petaFLOPMegatron-LM (1.2B)|2019|1.1×10^6 petaFLOPMegatron-LM (8.3B)|2019|9.1×10^6 petaFLOPAlphaX-1|2019|8.9×10^2 petaFLOPDistilBERT|2019|1.2×10^4 petaFLOPT5-11B|2019|3.3×10^7 petaFLOPT5-3B|2019|9.0×10^6 petaFLOPAlphaStar|2019|1.1×10^8 petaFLOPBase LM + kNN LM + Continuous Cache|2019|3.1×10^4 petaFLOPXLM-RoBERTa|2019|2.1×10^7 petaFLOPCamemBERT|2019|8.3×10^5 petaFLOPNoisy Student (L2)|2019|2.6×10^7 petaFLOPSandwich Transformer|2019|2.4×10^4 petaFLOPMuZero|2019|4.8×10^4 petaFLOPTransformer-XL DeFINE (141M)|2019|1.7×10^3 petaFLOPMMLSTM (PTB)|2019|5.8×10^1 petaFLOPMMLSTM (WT-2)|2019|1.9×10^2 petaFLOPOpenAI Five|2019|6.7×10^7 petaFLOPOpenAI Five Rerun|2019|1.3×10^7 petaFLOPDD-PPO|2019|7.8×10^5 petaFLOPAlphaFold|2020|1.0×10^5 petaFLOPContextNet + Noisy Student|2020|8.2×10^6 petaFLOPMeena|2020|1.1×10^8 petaFLOPTaLK Convolution|2020|2.7×10^4 petaFLOPALBERT-xxlarge|2020|2.4×10^6 petaFLOPFFN SwiGLU|2020|3.4×10^4 petaFLOPTuring-NLG|2020|1.6×10^7 petaFLOPFeedback Transformer|2020|7.7×10^3 petaFLOPTransformerXL + spectrum control|2020|2.6×10^4 petaFLOPTensor-Transformer(1core)+PN (WT103)|2020|1.6×10^3 petaFLOPELECTRA|2020|3.1×10^6 petaFLOPMetNet|2020|9.5×10^3 petaFLOPOnce for All|2020|6.2×10^5 petaFLOPUnifiedQA|2020|1.7×10^4 petaFLOPDETR|2020|4.0×10^5 petaFLOPGPT-3 175B (davinci)|2020|3.1×10^8 petaFLOPGShard (dense)|2020|4.8×10^7 petaFLOPDeLighT|2020|3.8×10^3 petaFLOPERNIE-GEN (large)|2020|2.0×10^5 petaFLOPProBERTa|2020|9.7×10^3 petaFLOPLUKE|2020|1.8×10^7 petaFLOPConformer + Wav2vec 2.0 + Noisy Student|2020|7.6×10^6 petaFLOPGerman ELECTRA Large|2020|1.4×10^6 petaFLOPmT5-XXL|2020|8.2×10^7 petaFLOPViT-Huge/14|2020|4.3×10^6 petaFLOPwave2vec 2.0 LARGE|2020|3.9×10^6 petaFLOPKEPLER|2020|1.7×10^6 petaFLOPAlphaFold 2|2020|3.0×10^6 petaFLOPCPM-Large|2020|2.6×10^5 petaFLOPESM1b|2020|5.1×10^6 petaFLOPCT-MoS (WT2)|2020|5.4×10^2 petaFLOPDensePhrases|2020|2.1×10^3 petaFLOPERNIE-Doc (247M)|2021|3.0×10^4 petaFLOPCLIP (ViT L/14@336px)|2021|1.1×10^7 petaFLOPDALL-E|2021|4.7×10^7 petaFLOPSwitch|2021|8.2×10^7 petaFLOPDeiT-B|2021|7.9×10^4 petaFLOPDLWP|2021|5.7×10^3 petaFLOPMSA Transformer|2021|5.5×10^6 petaFLOPSRU++ Large|2021|2.1×10^4 petaFLOPMeta Pseudo Labels|2021|4.8×10^7 petaFLOPGenerative BST|2021|1.4×10^7 petaFLOPM6-T|2021|5.5×10^6 petaFLOPPLUG|2021|3.6×10^7 petaFLOPProtBERT-BFD|2021|3.9×10^7 petaFLOPProtT5-XL-U50|2021|1.9×10^7 petaFLOPADM|2021|6.2×10^6 petaFLOPMedBERT|2021|9.5×10^3 petaFLOPByT5-XXL|2021|8.1×10^7 petaFLOPCogView|2021|2.7×10^7 petaFLOPTransformer local-attention (NesT-B)|2021|2.4×10^4 petaFLOPViT-G/14|2021|5.8×10^7 petaFLOPALIGN|2021|2.6×10^7 petaFLOPCoAtNet|2021|4.3×10^7 petaFLOPDeBERTa|2021|2.6×10^7 petaFLOPDenoising Diffusion Probabilistic Models (LSUN Bedroom)|2021|7.8×10^4 petaFLOPEMDR|2021|1.9×10^6 petaFLOPEfficientNetV2-XL|2021|9.6×10^4 petaFLOPStyleGAN3-R|2021|2.4×10^6 petaFLOPStyleGAN3-T|2021|1.7×10^6 petaFLOPFold2Seq|2021|1.4×10^2 petaFLOPAdaptive Input Transformer + RD|2021|8.6×10^4 petaFLOPCodex|2021|7.3×10^7 petaFLOPERNIE 3.0|2021|2.3×10^7 petaFLOPGOAT|2021|2.4×10^7 petaFLOPHuBERT|2021|5.5×10^6 petaFLOPSEER|2021|1.8×10^7 petaFLOPYOLOX-X|2021|6.3×10^5 petaFLOPJurassic-1-Jumbo|2021|3.7×10^8 petaFLOPZidong Taichu|2021|8.0×10^5 petaFLOPDNABERT|2021|1.1×10^5 petaFLOPXLMR-XXL|2021|3.4×10^7 petaFLOPFLAN 137B|2021|2.1×10^9 petaFLOPPermuteFormer|2021|2.8×10^3 petaFLOPHyperCLOVA 204B|2021|2.0×10^8 petaFLOPPLATO-XL|2021|9.9×10^6 petaFLOPTuring ULRv5|2021|2.9×10^7 petaFLOPAlphaFold-Multimer|2021|4.4×10^6 petaFLOPMegatron-Turing NLG 530B|2021|8.6×10^8 petaFLOPYuan 1.0|2021|3.5×10^8 petaFLOPbase LM+GNN+kNN|2021|5.3×10^4 petaFLOPCodeT5-base|2021|1.6×10^6 petaFLOPProjected GAN|2021|1.1×10^4 petaFLOPS4|2021|7.8×10^4 petaFLOPMasked Autoencoders ViT-H|2021|4.6×10^5 petaFLOPBASIC-L|2021|4.1×10^7 petaFLOPSwin Transformer V2 (SwinV2-G)|2021|1.1×10^6 petaFLOPFlorence|2021|4.8×10^7 petaFLOPNÜWA|2021|7.3×10^6 petaFLOPGopher (280B)|2021|6.3×10^8 petaFLOPStudent of Games|2021|3.7×10^7 petaFLOPGLaM|2021|3.6×10^8 petaFLOPContriever|2021|1.6×10^5 petaFLOPXGLM-7.5B|2021|2.3×10^7 petaFLOPERNIE 3.0 Titan|2021|1.0×10^9 petaFLOPDetic|2022|2.3×10^4 petaFLOPInstructGPT 175B|2022|3.2×10^8 petaFLOPAlphaCode|2022|2.4×10^8 petaFLOPRETRO-7B|2022|1.7×10^7 petaFLOPGPT-NeoX-20B|2022|9.3×10^7 petaFLOPLaMDA|2022|3.6×10^8 petaFLOPProteinBERT|2022|6.5×10^4 petaFLOPST-MoE|2022|2.9×10^8 petaFLOPFourCastNet|2022|3.5×10^5 petaFLOPPolyCoder|2022|1.1×10^6 petaFLOPStatement Curriculum Learning|2022|1.8×10^7 petaFLOPViT-G (model soup)|2022|3.4×10^6 petaFLOPGPT-3.5 (davinci-002)|2022|2.6×10^9 petaFLOPMake-A-Scene|2022|6.4×10^6 petaFLOPSegatron-XL large, M=384 + HCP|2022|2.7×10^4 petaFLOPChinchilla|2022|5.8×10^8 petaFLOPPaLM (540B)|2022|2.5×10^9 petaFLOPBERT-RBP|2022|1.4×10^5 petaFLOPDALL·E 2|2022|3.4×10^8 petaFLOPSparse all-MLP|2022|5.3×10^5 petaFLOPStable Diffusion (LDM-KL-8-G)|2022|5.0×10^7 petaFLOPFlamingo|2022|2.2×10^8 petaFLOPOPT-175B|2022|4.3×10^8 petaFLOPUL2|2022|1.2×10^8 petaFLOPGato|2022|4.0×10^6 petaFLOPImagen|2022|1.5×10^7 petaFLOPGPT-2 Medium (FlashAttention)|2022|8.9×10^5 petaFLOPTranception|2022|7.2×10^6 petaFLOPDITTO|2022|3.3×10^3 petaFLOPCoCa|2022|7.3×10^7 petaFLOPParti|2022|5.1×10^8 petaFLOPProGen2-xlarge|2022|1.3×10^7 petaFLOPMinerva (540B)|2022|2.7×10^9 petaFLOPCodeT5-large|2022|2.7×10^6 petaFLOPNLLB|2022|1.8×10^7 petaFLOPBLOOM-176B|2022|3.7×10^8 petaFLOPESM2-15B|2022|7.4×10^7 petaFLOPOmegaPLM|2022|1.0×10^7 petaFLOPAlexaTM 20B|2022|2.0×10^8 petaFLOPGLM-130B|2022|3.6×10^8 petaFLOPBlenderBot 3|2022|4.3×10^8 petaFLOPBEIT-3|2022|7.0×10^4 petaFLOPPaLI|2022|1.7×10^8 petaFLOPWhisper|2022|4.2×10^6 petaFLOPAlphaTensor|2022|7.1×10^5 petaFLOPDiffDock|2022|7.2×10^4 petaFLOPGenSLM|2022|1.4×10^6 petaFLOPFlan-PaLM 540B|2022|2.5×10^9 petaFLOPU-PaLM (540B)|2022|2.5×10^9 petaFLOPMogrifier RLSTM (WT2)|2022|1.4×10^2 petaFLOPeDiff-I|2022|5.5×10^4 petaFLOPInternImage|2022|2.4×10^6 petaFLOPEVA-01|2022|1.5×10^7 petaFLOPGalactica|2022|3.2×10^8 petaFLOPAR-LDM|2022|5.1×10^5 petaFLOPFusion in Encoder|2022|1.3×10^5 petaFLOPDiscriminator Guidance|2022|2.2×10^5 petaFLOPVega v2|2022|7.8×10^7 petaFLOPCaLM|2022|2.9×10^4 petaFLOPHybrid H3-2.7B|2022|6.5×10^6 petaFLOPVALL-E|2023|1.0×10^4 petaFLOPDreamerV3|2023|2.2×10^5 petaFLOPAnkh_large|2023|6.5×10^6 petaFLOPNucleotide Transformer|2023|8.1×10^6 petaFLOPDDPM-IP (CelebA)|2023|3.5×10^5 petaFLOPBLIP-2 (Q-Former)|2023|1.2×10^6 petaFLOPViT-22B|2023|1.9×10^8 petaFLOPLLaMA-65B|2023|5.5×10^8 petaFLOPDiT-XL/2|2023|6.0×10^5 petaFLOPAudioGen|2023|9.5×10^6 petaFLOPFalcon-40B|2023|2.4×10^8 petaFLOPGPT-4 (Mar 2023)|2023|2.1×10^10 petaFLOPPanGu-Σ|2023|4.7×10^8 petaFLOPSigLIP 400M|2023|5.0×10^6 petaFLOPBloombergGPT|2023|2.4×10^8 petaFLOPVideoMAE V2|2023|9.7×10^6 petaFLOPSegment Anything Model|2023|7.8×10^6 petaFLOPIncoder-6.7B|2023|3.0×10^6 petaFLOPDINOv2|2023|7.4×10^6 petaFLOPLLaVA|2023|7.8×10^7 petaFLOPPaLM 2|2023|7.3×10^9 petaFLOPStarCoder|2023|8.5×10^7 petaFLOPInstructBLIP|2023|1.9×10^5 petaFLOPONE-PEACE|2023|1.8×10^5 petaFLOPPaLI-X|2023|5.6×10^8 petaFLOPGPT-4 (Jun 2023)|2023|2.1×10^10 petaFLOPHyenaDNA|2023|1.8×10^6 petaFLOPInternLM|2023|1.0×10^9 petaFLOPPangu-Weather|2023|4.0×10^7 petaFLOPxTrimoPGLM -100B|2023|6.2×10^8 petaFLOPClaude 2|2023|3.9×10^9 petaFLOPLlama 2-70B|2023|8.1×10^8 petaFLOPLlama 2-7B|2023|8.4×10^7 petaFLOPAudioLM|2023|3.9×10^3 petaFLOPGGNN|2023|7.6×10^6 petaFLOPPeptideBERT|2023|4.9×10^1 petaFLOPJais|2023|4.9×10^7 petaFLOPSwift|2023|5.3×10^1 petaFLOPFalcon-180B|2023|3.8×10^9 petaFLOPAmazon Titan|2023|4.8×10^9 petaFLOPFinGPT-13B|2023|1.6×10^8 petaFLOPRoseTTAFold All-Atom (RFAA)|2023|2.1×10^5 petaFLOPCODEFUSION (Python)|2023|7.9×10^3 petaFLOPChatGLM3-6B|2023|5.0×10^7 petaFLOPSkywork-13B|2023|2.5×10^8 petaFLOPGrok-1|2023|2.9×10^9 petaFLOPLLaVA 1.5|2023|7.8×10^7 petaFLOPYi-34B|2023|6.1×10^8 petaFLOPCogVLM-17B|2023|6.3×10^7 petaFLOPMultiBand Diffusion|2023|2.6×10^4 petaFLOPRoFormer|2023|2.2×10^3 petaFLOPGraphCast|2023|2.1×10^7 petaFLOPNemotron-3-8B|2023|1.8×10^8 petaFLOPSPHINX (Llama 2 13B)|2023|3.0×10^7 petaFLOPVolcano 13B|2023|4.6×10^7 petaFLOPInflection-2|2023|1.0×10^10 petaFLOPQwen-72B|2023|1.3×10^9 petaFLOPGemini 1.0 Ultra|2023|5.0×10^10 petaFLOPLlama Guard|2023|1.6×10^8 petaFLOPMixtral 8x7B|2023|7.7×10^8 petaFLOPCogAgent|2023|6.7×10^7 petaFLOPFunSearch|2023|3.9×10^8 petaFLOPVILA-13B|2023|2.3×10^6 petaFLOPnekomata-14b|2023|2.6×10^8 petaFLOPGQA-8-XXL|2023|3.5×10^7 petaFLOPQwen1.5-72B|2024|1.3×10^9 petaFLOPMegaScale (Production)|2024|3.9×10^9 petaFLOPStable Diffusion 3|2024|5.0×10^7 petaFLOPAramco Metabrain AI|2024|1.1×10^10 petaFLOPInflection-2.5|2024|8.0×10^9 petaFLOPMM1-30B|2024|4.9×10^8 petaFLOPDBRX|2024|2.6×10^9 petaFLOPReka Core|2024|8.4×10^9 petaFLOPLlama 3-70B|2024|7.9×10^9 petaFLOPGenCast|2024|8.2×10^5 petaFLOPVILA1.5-13B|2024|2.3×10^6 petaFLOPAlphaFold 3|2024|4.1×10^7 petaFLOPYi-Large|2024|1.8×10^9 petaFLOPOcto-Base|2024|5.9×10^5 petaFLOPALLaM adapted 70B|2024|1.1×10^9 petaFLOPQwen2-72B|2024|3.0×10^9 petaFLOPNemotron-4 340B|2024|1.8×10^10 petaFLOPOpenVLA|2024|1.1×10^8 petaFLOPDeepSeek-Coder-V2 236B|2024|1.3×10^9 petaFLOPClaude 3.5 Sonnet|2024|2.7×10^10 petaFLOPESM3 (98B)|2024|1.1×10^9 petaFLOPLlama 3.1-405B|2024|3.8×10^10 petaFLOPAFM-on-device|2024|4.5×10^8 petaFLOPAFM-server|2024|4.3×10^9 petaFLOPLLaVA-OV-72B|2024|3.0×10^9 petaFLOPGrok-2|2024|3.0×10^10 petaFLOPDeepSeek-V2.5|2024|1.8×10^9 petaFLOPQwen2.5-32B|2024|3.5×10^9 petaFLOPQwen2.5 Instruct (72B)|2024|7.9×10^9 petaFLOPQwen2.5-72B|2024|7.8×10^9 petaFLOPTelechat2-115B|2024|6.9×10^9 petaFLOPLlama 3.2 11B|2024|5.8×10^8 petaFLOPMovie Gen Video|2024|1.7×10^9 petaFLOPRDT-1B|2024|4.1×10^7 petaFLOPLlama-3.1-Nemotron-70B-Instruct|2024|7.9×10^9 petaFLOPCHAI-1|2024|7.8×10^6 petaFLOPYi-Lightning|2024|1.5×10^9 petaFLOPNVLM-D 72B|2024|3.0×10^9 petaFLOPNVLM-H 72B|2024|3.0×10^9 petaFLOPNVLM-X 72B|2024|3.0×10^9 petaFLOPDoubao-pro|2024|2.5×10^10 petaFLOPHunyuan-Large|2024|3.5×10^9 petaFLOPAmazon Nova Pro|2024|6.0×10^9 petaFLOPLlama 3.3 70B|2024|6.9×10^9 petaFLOPEXAONE 3.5 32B|2024|1.3×10^9 petaFLOPDeepSeek-V3|2024|3.3×10^9 petaFLOPDeepSeek-R1|2025|3.5×10^9 petaFLOPEagle 2|2025|4.7×10^7 petaFLOPEurus-2-7B-PRIME|2025|4.0×10^5 petaFLOPGrok 3|2025|3.5×10^11 petaFLOPClaude 3.7 Sonnet|2025|3.4×10^10 petaFLOPGPT-4.5|2025|3.8×10^11 petaFLOPQwQ-32B|2025|3.5×10^9 petaFLOPHunyuan-TurboS|2025|5.4×10^9 petaFLOPEXAONE Deep 32B|2025|1.3×10^9 petaFLOPDeepSeek-V3 (Mar 2025)|2025|3.3×10^9 petaFLOPLlama 4 Behemoth (preview)|2025|5.2×10^10 petaFLOPLlama 4 Maverick|2025|2.2×10^9 petaFLOPLlama 4 Scout|2025|4.1×10^9 petaFLOPPangu Ultra|2025|1.1×10^10 petaFLOPQwen3-235B-A22B|2025|4.8×10^9 petaFLOPSeed1.5-VL|2025|1.4×10^9 petaFLOPDeepSeek-R1 (May 2025)|2025|4.0×10^9 petaFLOPFGN|2025|9.6×10^6 petaFLOPGrok 4|2025|5.0×10^11 petaFLOPKimi K2|2025|3.0×10^9 petaFLOPEXAONE 4.0 (32B)|2025|2.7×10^9 petaFLOPQwen3-Coder-480B-A35B|2025|1.6×10^9 petaFLOPQwen3-235B-A22B (Jul 2025)|2025|4.8×10^9 petaFLOPQwen3-235B-A22B-Thinking (Jul 2025)|2025|4.8×10^9 petaFLOPGLM-4.5|2025|4.4×10^9 petaFLOPgpt-oss-120b|2025|4.9×10^9 petaFLOPgpt-oss-20b|2025|5.5×10^8 petaFLOPGPT-5|2025|6.6×10^10 petaFLOPLongCat-Flash|2025|3.7×10^9 petaFLOPQwen3-Max|2025|1.5×10^10 petaFLOPAgentFounder-30B|2025|6.5×10^8 petaFLOPQwen3-Omni-30B-A3B|2025|3.6×10^7 petaFLOPGLM-4.6|2025|4.4×10^9 petaFLOPLing-1T|2025|6.0×10^9 petaFLOPKimi K2 Thinking|2025|4.2×10^9 petaFLOPOlmo 3|2025|1.1×10^9 petaFLOPNemotron 3-Nano-30B-A3B|2025|4.8×10^8 petaFLOPGLM-4.7|2025|4.4×10^9 petaFLOPK-EXAONE|2026|1.5×10^9 petaFLOPSolar Open 100B|2026|1.4×10^9 petaFLOPKimi K2.5|2026|5.8×10^9 petaFLOPGLM-5|2026|6.8×10^9 petaFLOPComposer 2|2026|2.3×10^10 petaFLOPMolmoAct 2|2026|1.1×10^7 petaFLOPEXAONE 4.5|2026|3.9×10^9 petaFLOPDeepSeek-V4-Flash|2026|2.5×10^9 petaFLOPDeepSeek-V4-Pro|2026|9.7×10^9 petaFLOPMiMo-V2.5-Pro|2026|6.8×10^9 petaFLOPComposer 2.5|2026|3.9×10^10 petaFLOPNemotron 3 Ultra|2026|6.6×10^9 petaFLOPSolar Open2 250B|2026|1.1×10^9 petaFLOPInkling|2026|1.9×10^9 petaFLOPKimi K3|2026|2.0×10^10 petaFLOPA.X K2|2026|1.8×10^9 petaFLOPK-EXAONE 2.0|2026|3.6×10^9 petaFLOPMotif-3|2026|1.1×10^9 petaFLOP其他
每個點為一個知名 AI 模型,橫軸=釋出年份、縱軸=訓練算力(對數軸,單位 petaFLOP=10¹⁵ 次浮點運算),按應用領域著色;共 535 個模型。

引數量趨勢

模型的可訓練引數數量——引數越多,模型「容量」通常越大。按開發方型別(產業界 / 學術界 / 產學合作)著色,可見近年前沿被產業界主導。

10²10⁴10⁶10⁸10¹⁰10¹²1960198020002020釋出年份引數量(個,對數軸)Theseus|1950|4.0×10^1 個引數SNARC|1952|4.0×10^1 個引數Self Organizing System|1955|2.3×10^2 個引數Perceptron Mark I|1957|1.0×10^3 個引數Samuel Neural Checkers|1959|1.6×10^1 個引數Pattern recognition and reading by machine|1959|2.6×10^3 個引數Perceptron (1960)|1960|1.0×10^3 個引數ADALINE|1960|1.7×10^1 個引數LTE speaker verification system|1966|2.1×10^3 個引數Decision tree adaline|1969|2.5×10^3 個引數Piecewise linear model|1973|3.6×10^2 個引數Cognitron|1975|2.2×10^4 個引數Neocognitron|1980|1.1×10^6 個引數Kohonen network|1981|4.1×10^3 個引數Hopfield network|1982|9.9×10^3 個引數ASE+ACE|1983|3.2×10^2 個引數Hierarchical Cognitron|1984|9.3×10^3 個引數Distributed representation NN|1986|4.3×10^2 個引數MLP with back-propagation|1986|7.2×10^2 個引數NetTalk (dictionary)|1987|1.9×10^4 個引數NetTalk (transcription)|1987|1.9×10^4 個引數Translation-invariant MLP|1987|8.2×10^2 個引數MLN-ASR|1988|1.0×10^4 個引數Truck backer-upper|1989|8.1×10^2 個引數Handwritten digit recognition network|1989|2.6×10^3 個引數Speaker-independent vowel classification|1989|3.0×10^3 個引數Zip CNN|1989|9.8×10^3 個引數NETtalk reimplementation|1990|2.8×10^4 個引數Bankruptcy-NN|1990|3.6×10^1 個引數SexNet classification|1990|1.6×10^3 個引數SexNet compression|1990|7.3×10^4 個引數RAAM|1990|1.5×10^3 個引數Weight Decay|1991|8.4×10^3 個引數TD-Gammon|1992|2.5×10^4 個引數Cancer drug mechanism prediction|1992|5.9×10^2 個引數Boosting|1992|2.6×10^3 個引數IBM-5|1993|1.7×10^6 個引數Siamese-TDNN|1993|7.4×10^2 個引數ANN Eye Tracker|1993|5.6×10^3 個引數Ceramic-MLP|1994|1.9×10^3 個引數JPMAX|1994|4.5×10^3 個引數Mixture of linear models|1994|3.8×10^5 個引數NeuroChess|1994|7.2×10^4 個引數Predictive Coding NN|1994|2.1×10^5 個引數Support Vector Machines|1995|1.0×10^8 個引數LISSOM|1995|4.3×10^5 個引數MUSIC perceptron|1996|1.4×10^4 個引數System 11|1996|6.5×10^3 個引數SOM-CNN|1997|3.2×10^4 個引數Deep Blue|1997|8.0×10^3 個引數Bidirectional RNN|1997|1.3×10^4 個引數LSTM|1997|1.1×10^4 個引數LeNet-5|1998|6.0×10^4 個引數LSTM with forget gates|1999|2.8×10^2 個引數RECONTRA-categorized|1999|6.7×10^4 個引數RECONTRA-uncategorized|1999|1.1×10^5 個引數Neural LM|2000|6.9×10^6 個引數PoE MNIST|2000|3.9×10^6 個引數Decision tree (classification)|2001|1.2×10^4 個引數NPLM (AP News)|2003|1.2×10^7 個引數NPLM (Brown)|2003|4.1×10^6 個引數Invariant CNN|2004|9.1×10^4 個引數LMICA|2004|4.1×10^6 個引數RankNet|2005|5.7×10^3 個引數SVM-CNN|2006|9.1×10^4 個引數Deep Belief Nets|2006|1.6×10^6 個引數Dimensionality Reduction|2006|3.8×10^6 個引數KN-LM|2007|2.1×10^10 個引數SB-LM|2007|3.0×10^11 個引數BLSTM for handwriting (2)|2007|1.0×10^5 個引數Deep Multitask NLP Network|2008|1.5×10^6 個引數HLBL|2008|1.9×10^6 個引數GNN|2008|3.0×10^1 個引數BP-DBN|2009|1.8×10^7 個引數RBM Image Classifier|2009|8.0×10^7 個引數GPU DBNs|2009|1.0×10^8 個引數Two Stage Feature Extraction (MNIST)|2009|2.6×10^5 個引數LCNP LabelMe|2009|1.4×10^7 個引數LCNP MNIST|2009|1.2×10^7 個引數LCNP NORB|2009|1.7×10^7 個引數Super-vector coding|2010|1.0×10^3 個引數Feedforward NN|2010|7.1×10^6 個引數ReLU (NORB)|2010|1.6×10^7 個引數Pooling CNN (Caltech 101)|2010|3.0×10^5 個引數Pooling CNN (NORB)|2010|2.7×10^5 個引數RNN LM|2010|7.0×10^7 個引數Deep Autoencoders|2011|1.4×10^8 個引數Vector Space Model|2011|2.6×10^5 個引數High Performance CNN (NORB)|2011|4.9×10^6 個引數CNN Committee (MNIST)|2011|1.2×10^5 個引數CNN Committee (NIST)|2011|1.3×10^5 個引數CNN committee (traffic sign)|2011|1.4×10^6 個引數NLP from scratch|2011|5.0×10^6 個引數Dropout (MNIST)|2012|5.6×10^6 個引數Dropout (TIMIT)|2012|4.9×10^7 個引數Unsupervised High-level Feature Learner|2012|1.0×10^9 個引數LSTM LM|2012|1.0×10^8 個引數AlexNet|2012|6.0×10^7 個引數DNN EM segmentation|2012|2.2×10^5 個引數DistBelief Speech|2012|4.7×10^7 個引數DistBelief Vision|2012|1.7×10^9 個引數RNN+LDA+KN5+cache|2012|9.0×10^6 個引數PreTrans-3L-250H|2013|4.3×10^7 個引數Multilingual DNN|2013|2.1×10^8 個引數ReLU-Speech|2013|1.0×10^8 個引數Hierarchical Scene Labeling (Stanford Background)|2013|5.2×10^7 個引數Word2Vec (large)|2013|6.9×10^8 個引數Word2Vec (small)|2013|2.1×10^8 個引數R-CNN (T-net)|2013|6.9×10^7 個引數TransE|2013|9.4×10^8 個引數RNN for 1B words|2013|2.0×10^10 個引數DQN|2013|8.4×10^5 個引數Image generation|2013|7.8×10^5 個引數OverFeat|2013|1.4×10^8 個引數GloVe (32B)|2014|1.2×10^8 個引數GloVe (6B)|2014|1.2×10^8 個引數HyperNEAT|2014|2.4×10^5 個引數Paragraph Vector|2014|3.2×10^7 個引數AdaRNN|2014|1.3×10^4 個引數Dropout: SVHN|2014|4.8×10^7 個引數Fragment embedding|2014|1.4×10^8 個引數Multiresolution CNN|2014|1.3×10^8 個引數RNN-WER|2014|2.6×10^7 個引數ACF-WIDER|2014|6.1×10^3 個引數NPD|2014|3.1×10^5 個引數VGG16|2014|1.4×10^8 個引數VGG19|2014|1.4×10^8 個引數Seq2Seq LSTM|2014|1.9×10^9 個引數SPN-4+KN5|2014|5.0×10^6 個引數GoogLeNet / InceptionV1|2014|6.8×10^6 個引數LRCN|2014|1.4×10^8 個引數TA-CNN|2014|7.1×10^5 個引數SNM-skip|2014|6.2×10^10 個引數Fractional Max-Pooling|2014|2.7×10^7 個引數ADAM (CIFAR-10)|2014|2.4×10^6 個引數VGG-Face|2015|1.4×10^8 個引數MSRA (C, PReLU)|2015|8.7×10^7 個引數TRPO|2015|3.4×10^4 個引數DQN-2015|2015|1.7×10^6 個引數genCNN + dyn eval|2015|8.0×10^6 個引數TC-DNN-BLSTM-DNN|2015|1.8×10^7 個引數U-Net|2015|3.8×10^7 個引數CFSS|2015|1.7×10^4 個引數YOLO|2015|2.7×10^8 個引數BatchNorm|2015|1.4×10^7 個引數Deep CNN + COTS|2015|5.0×10^6 個引數DCNN|2015|5.0×10^6 個引數AlphaGo Fan|2015|8.2×10^6 個引數SAF R-CNN|2015|1.4×10^8 個引數3DDFA|2015|5.4×10^6 個引數Inception v3|2015|2.4×10^7 個引數ResNet-101 (ImageNet)|2015|4.5×10^7 個引數ResNet-110 (CIFAR-10)|2015|1.7×10^6 個引數ResNet-152 (ImageNet)|2015|6.0×10^7 個引數Variational (untied weights, MC) LSTM (Large)|2015|6.6×10^7 個引數Inception-ResNet-V2|2016|5.6×10^7 個引數Inceptionv4|2016|4.3×10^7 個引數SqueezeNet|2016|1.2×10^6 個引數Double DQN|2016|1.5×10^6 個引數Template Adaptation|2016|1.4×10^8 個引數Dueling DQN|2016|1.7×10^6 個引數Gated HORNN (3rd order)|2016|9.0×10^6 個引數LRR-4X|2016|1.4×10^8 個引數CMS-RCNN|2016|1.4×10^8 個引數SimpleNet|2016|5.5×10^6 個引數DenseNet-264|2016|3.4×10^7 個引數LF-MMI|2016|1.7×10^7 個引數MS-ensemble-speech-recognition|2016|3.2×10^9 個引數ResNet-1001|2016|1.0×10^7 個引數GNMT|2016|2.8×10^8 個引數Pointer Sentinel-LSTM (medium)|2016|2.1×10^7 個引數Xception|2016|2.3×10^7 個引數SPIDER2|2016|4.1×10^5 個引數BIDAF|2016|2.6×10^6 個引數NAS with base 8 and shared embeddings|2016|5.4×10^7 個引數NASv3 (CIFAR-10)|2016|3.7×10^7 個引數VD-LSTM+REAL Large|2016|5.1×10^7 個引數DLDL (PASCAL)|2016|5.6×10^8 個引數ResNeXt-101 (64×4d)|2016|8.3×10^7 個引數ResNeXt-50|2016|2.5×10^7 個引數PolyNet|2016|9.2×10^7 個引數3DMM-CNN|2016|4.5×10^7 個引數HR-ResNet101|2016|4.5×10^7 個引數EnhanceNet|2016|8.1×10^5 個引數YOLOv2|2016|5.1×10^7 個引數DeepStack|2017|2.5×10^6 個引數OR-WideResNet|2017|1.8×10^7 個引數MoE-Multi|2017|8.7×10^9 個引數MobileNet|2017|4.2×10^6 個引數Transformer (2017)|2017|2.1×10^8 個引數ShuffleNet v1|2017|2.4×10^6 個引數JFT|2017|4.5×10^7 個引數AWD-LSTM|2017|2.4×10^7 個引數NASNet-A|2017|8.9×10^7 個引數AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2)|2017|3.3×10^7 個引數RetinaNet-R101|2017|5.3×10^7 個引數RetinaNet-R50|2017|3.4×10^7 個引數EI-REHN-1000D|2017|1.9×10^7 個引數GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2)|2017|3.8×10^7 個引數PyramidNet|2017|2.6×10^7 個引數SENet (ImageNet)|2017|2.8×10^7 個引數ISS|2017|1.1×10^7 個引數LSTM + dynamic eval|2017|5.0×10^7 個引數AWD-LSTM+WT+Cache+IOG (WT2)|2017|5.3×10^7 個引數AlphaGo Zero|2017|4.6×10^7 個引數Fraternal dropout + AWD-LSTM 3-layer (WT2)|2017|3.4×10^7 個引數AWD-LSTM-MoS + dynamic evaluation (WT2, 2017)|2017|3.5×10^7 個引數DL scaling Image|2017|1.2×10^8 個引數DL scaling LM|2017|1.8×10^8 個引數DL scaling speech|2017|1.9×10^8 個引數ELMo|2018|9.4×10^7 個引數QRNN|2018|1.4×10^8 個引數IMPALA|2018|1.6×10^6 個引數TCN (P-MNIST)|2018|4.2×10^4 個引數4 layer QRNN (h=2500)|2018|1.5×10^8 個引數YOLOv3|2018|5.7×10^7 個引數Dropout-LSTM+Noise(Bernoulli) (WT2)|2018|5.1×10^7 個引數ResNeXt-101 32x48d|2018|8.3×10^8 個引數aLSTM(depth-2)+RecurrentPolicy (WT2)|2018|3.2×10^7 個引數GPT-1|2018|1.2×10^8 個引數MobileNetV2|2018|3.4×10^6 個引數FTW (For The Win)|2018|1.3×10^8 個引數Big-Little Net|2018|7.7×10^7 個引數Big-Little Net (speech)|2018|3.3×10^6 個引數AWD-LSTM-MoS+PDR + dynamic evaluation (WT2)|2018|3.5×10^7 個引數(ensemble): AWD-LSTM-DOC (fin) × 5 (WT2)|2018|1.9×10^8 個引數AWD-LSTM-MoS + dynamic evaluation (WT2, 2018)|2018|3.5×10^7 個引數Transformer + Simple Recurrent Unit|2018|9.0×10^7 個引數LSTM+NeuralCache|2018|2.1×10^6 個引數Transformer (Adaptive Input Embeddings) WT103|2018|2.5×10^8 個引數BERT-Large|2018|3.4×10^8 個引數MetaMimic|2018|2.2×10^7 個引數TrellisNet|2018|1.8×10^8 個引數Mesh-TensorFlow Transformer 2.9B (translation)|2018|2.9×10^9 個引數Mesh-TensorFlow Transformer 4.9B (language)|2018|4.9×10^9 個引數Fine-tuned-AWD-LSTM-DOC (fin)|2018|4.6×10^7 個引數GPipe (Transformer)|2018|6.0×10^9 個引數Multi-cell LSTM|2018|7.2×10^6 個引數SPN (ImageNet 128)|2018|2.5×10^8 個引數StyleGAN|2018|2.6×10^7 個引數Transformer ELMo|2019|5.6×10^7 個引數Transformer-XL (257M)|2019|2.6×10^8 個引數Hanabi 4 player|2019|7.6×10^5 個引數MT-DNN|2019|3.3×10^8 個引數GPT-2 (1.5B)|2019|1.5×10^9 個引數KataGo|2019|2.5×10^6 個引數NMT Transformer 437M|2019|4.4×10^8 個引數SciBERT|2019|1.1×10^8 個引數True-Regularization+Finetune+Dynamic-Eval|2019|7.0×10^6 個引數WeNet (Penn Treebank)|2019|2.3×10^7 個引數Transformer-XL + RMS dynamic eval|2019|2.6×10^8 個引數BERT-Large-CAS (PTB+WT2+WT103)|2019|4.0×10^8 個引數MuseNet|2019|2.0×10^9 個引數ResNeXt-101 Billion-scale|2019|1.9×10^8 個引數AWD-LSTM-DRILL + dynamic evaluation† (WT2)|2019|3.4×10^7 個引數CPC v2|2019|3.0×10^8 個引數EfficientNet-L2|2019|4.8×10^8 個引數DLRM-2020|2019|1.0×10^11 個引數XLM|2019|6.7×10^8 個引數XLNet|2019|3.4×10^8 個引數Transformer-XL Large + Phrase Induction|2019|2.6×10^8 個引數AWD-LSTM + MoS + Partial Shuffled|2019|3.5×10^7 個引數FixRes ResNeXt-101 WSL|2019|8.3×10^8 個引數LaNet-L (CIFAR-10)|2019|4.4×10^7 個引數BigBiGAN|2019|8.6×10^7 個引數RoBERTa Large|2019|3.6×10^8 個引數EN^2AS with performance reward|2019|2.3×10^7 個引數Mogrifier (d2, MoS2, MC) + dynamic eval|2019|3.5×10^7 個引數UDSMProt|2019|2.8×10^7 個引數Megatron-BERT|2019|3.9×10^9 個引數Megatron-LM (1.2B)|2019|1.2×10^9 個引數Megatron-LM (8.3B)|2019|8.3×10^9 個引數ALBERT|2019|1.8×10^7 個引數Adaptive Inputs + LayerDrop|2019|4.2×10^8 個引數AlphaX-1|2019|5.4×10^6 個引數DistilBERT|2019|6.6×10^7 個引數M4-50B|2019|5.0×10^10 個引數T5-11B|2019|1.1×10^10 個引數T5-3B|2019|2.8×10^9 個引數BART-large|2019|4.1×10^8 個引數AlphaStar|2019|1.4×10^8 個引數Base LM + kNN LM + Continuous Cache|2019|2.5×10^8 個引數XLM-RoBERTa|2019|5.5×10^8 個引數CamemBERT|2019|3.4×10^8 個引數Noisy Student (L2)|2019|4.8×10^8 個引數Sandwich Transformer|2019|2.1×10^8 個引數MoCo|2019|3.8×10^8 個引數MuZero|2019|3.7×10^7 個引數Transformer - LibriVox + Decoding/Rescoring|2019|3.0×10^8 個引數Transformer-XL DeFINE (141M)|2019|1.4×10^8 個引數StyleGAN2|2019|3.0×10^7 個引數MMLSTM (PTB)|2019|2.1×10^7 個引數MMLSTM (WT-2)|2019|3.2×10^7 個引數OpenAI Five|2019|1.6×10^8 個引數OpenAI Five Rerun|2019|1.6×10^8 個引數Big Transfer (BiT-L)|2019|9.3×10^8 個引數AlphaFold|2020|1.6×10^7 個引數Meena|2020|2.6×10^9 個引數Perceiver IO (optical flow)|2020|2.8×10^7 個引數TaLK Convolution|2020|2.4×10^8 個引數Theseus 6/768|2020|6.6×10^7 個引數ALBERT-xxlarge|2020|2.4×10^8 個引數FFN SwiGLU|2020|2.2×10^8 個引數SimCLR|2020|3.8×10^8 個引數Turing-NLG|2020|1.7×10^10 個引數Feedback Transformer|2020|1.3×10^8 個引數TCAN (WT2)|2020|3.3×10^7 個引數Routing Transformer (WT-103)|2020|8.0×10^7 個引數TransformerXL + spectrum control|2020|1.5×10^8 個引數Tensor-Transformer(1core)+PN (WT103)|2020|8.5×10^7 個引數ELECTRA|2020|3.4×10^8 個引數MetNet|2020|2.3×10^8 個引數CURL|2020|9.1×10^5 個引數Once for All|2020|7.7×10^6 個引數UnifiedQA|2020|1.1×10^10 個引數NAS+ESS (23M)|2020|2.3×10^7 個引數ContextNet|2020|1.1×10^8 個引數Conformer|2020|1.2×10^8 個引數Retrieval-Augmented Generator|2020|6.3×10^8 個引數DETR|2020|6.0×10^7 個引數GPT-3 175B (davinci)|2020|1.8×10^11 個引數GShard (dense)|2020|2.3×10^9 個引數EfficientDet|2020|7.7×10^7 個引數DeLighT|2020|9.9×10^7 個引數ERNIE-GEN (large)|2020|3.4×10^8 個引數ProBERTa|2020|4.4×10^7 個引數LUKE|2020|4.8×10^8 個引數Conformer + Wav2vec 2.0 + Noisy Student|2020|1.0×10^9 個引數German ELECTRA Large|2020|3.4×10^8 個引數mT5-XXL|2020|1.3×10^10 個引數ViT-Base/32|2020|8.6×10^7 個引數ViT-Huge/14|2020|6.3×10^8 個引數wave2vec 2.0 LARGE|2020|3.2×10^8 個引數KEPLER|2020|1.3×10^8 個引數AlphaFold 2|2020|9.3×10^7 個引數CPM-Large|2020|2.6×10^9 個引數ESM1b|2020|6.5×10^8 個引數CT-MoS (WT2)|2020|4.5×10^7 個引數ERNIE-Doc (247M)|2021|2.5×10^8 個引數CLIP (ResNet-50)|2021|8.9×10^7 個引數CLIP (ViT L/14@336px)|2021|3.7×10^8 個引數DALL-E|2021|1.2×10^10 個引數BigSSL|2021|8.0×10^9 個引數Switch|2021|1.6×10^12 個引數DeiT-B|2021|8.6×10^7 個引數DLWP|2021|2.7×10^6 個引數MSA Transformer|2021|1.0×10^8 個引數Rational DQN Average|2021|1.7×10^6 個引數SRU++ Large|2021|2.3×10^8 個引數Meta Pseudo Labels|2021|4.8×10^8 個引數Generative BST|2021|9.4×10^9 個引數M6-T|2021|1.0×10^12 個引數Unicorn|2021|1.1×10^10 個引數PLUG|2021|2.7×10^10 個引數ProtBERT-BFD|2021|4.2×10^8 個引數ProtT5-XL-U50|2021|3.0×10^9 個引數ADM|2021|5.6×10^8 個引數MedBERT|2021|1.7×10^7 個引數ByT5-XXL|2021|1.3×10^10 個引數CogView|2021|4.0×10^9 個引數Transformer local-attention (NesT-B)|2021|9.0×10^7 個引數ViT-G/14|2021|1.8×10^9 個引數ALIGN|2021|8.2×10^8 個引數CoAtNet|2021|2.4×10^9 個引數DeBERTa|2021|1.5×10^9 個引數Denoising Diffusion Probabilistic Models (LSUN Bedroom)|2021|2.6×10^8 個引數EMDR|2021|4.4×10^8 個引數EfficientNetV2-XL|2021|2.1×10^8 個引數StyleGAN3-R|2021|1.6×10^6 個引數StyleGAN3-T|2021|2.2×10^6 個引數Fold2Seq|2021|1.2×10^7 個引數Adaptive Input Transformer + RD|2021|2.5×10^8 個引數Codex|2021|1.2×10^10 個引數ERNIE 3.0|2021|1.0×10^10 個引數GOAT|2021|3.5×10^6 個引數HuBERT|2021|1.0×10^9 個引數SEER|2021|1.3×10^9 個引數6-Act Tether|2021|5.0×10^6 個引數YOLOX-X|2021|9.9×10^7 個引數W2v-BERT|2021|1.0×10^9 個引數Jurassic-1-Jumbo|2021|1.8×10^11 個引數Zidong Taichu|2021|3.2×10^9 個引數DNABERT|2021|1.1×10^8 個引數XLMR-XXL|2021|1.1×10^10 個引數FLAN 137B|2021|1.4×10^11 個引數MEB|2021|1.3×10^11 個引數PermuteFormer|2021|1.5×10^8 個引數HyperCLOVA 204B|2021|2.0×10^11 個引數PLATO-XL|2021|1.1×10^10 個引數TrOCR|2021|5.6×10^8 個引數Turing ULRv5|2021|2.2×10^9 個引數Megatron-Turing NLG 530B|2021|5.3×10^11 個引數Yuan 1.0|2021|2.5×10^11 個引數base LM+GNN+kNN|2021|2.7×10^8 個引數Eve|2021|1.5×10^7 個引數CodeT5-base|2021|2.2×10^8 個引數S4|2021|2.5×10^8 個引數Masked Autoencoders ViT-H|2021|6.3×10^8 個引數ViT-G/14 (LiT)|2021|3.0×10^9 個引數BASIC-L|2021|3.1×10^9 個引數Swin Transformer V2 (SwinV2-G)|2021|3.0×10^9 個引數Florence|2021|8.9×10^8 個引數NÜWA|2021|8.7×10^8 個引數T-NLRv5 XXL|2021|5.4×10^9 個引數Gopher (280B)|2021|2.8×10^11 個引數GLaM|2021|1.2×10^12 個引數LongT5|2021|3.0×10^9 個引數Contriever|2021|1.1×10^8 個引數LDM-1.45B|2021|1.5×10^9 個引數XGLM-7.5B|2021|7.5×10^9 個引數ERNIE 3.0 Titan|2021|2.6×10^11 個引數ERNIE-ViLG|2022|1.0×10^10 個引數Detic|2022|8.8×10^7 個引數data2vec (language)|2022|7.1×10^8 個引數data2vec (speech)|2022|7.1×10^8 個引數data2vec (vision)|2022|7.1×10^8 個引數AbLang (heavy sequences)|2022|3.6×10^8 個引數OntoProtein|2022|4.2×10^8 個引數InstructGPT 1.3B|2022|1.3×10^9 個引數InstructGPT 175B|2022|1.8×10^11 個引數InstructGPT 6B|2022|6.0×10^9 個引數AlphaCode|2022|4.1×10^10 個引數MaskGIT (ImageNet)|2022|2.3×10^8 個引數RETRO-7B|2022|7.5×10^9 個引數GPT-NeoX-20B|2022|2.0×10^10 個引數LaMDA|2022|1.4×10^11 個引數ProteinBERT|2022|1.6×10^7 個引數ST-MoE|2022|2.7×10^11 個引數PolyCoder|2022|2.7×10^9 個引數DeepNet|2022|3.2×10^9 個引數Statement Curriculum Learning|2022|7.7×10^8 個引數ViT-G (model soup)|2022|1.8×10^9 個引數Make-A-Scene|2022|4.0×10^9 個引數Segatron-XL large, M=384 + HCP|2022|2.6×10^8 個引數Chinchilla|2022|7.0×10^10 個引數PaLM (540B)|2022|5.4×10^11 個引數BERT-RBP|2022|1.1×10^8 個引數DALL·E 2|2022|3.5×10^9 個引數Sparse all-MLP|2022|9.4×10^9 個引數Stable Diffusion (LDM-KL-8-G)|2022|1.5×10^9 個引數Flamingo|2022|8.0×10^10 個引數OPT-175B|2022|1.8×10^11 個引數DeBERTaV3large + KEAR|2022|4.2×10^8 個引數UL2|2022|2.0×10^10 個引數Gato|2022|1.2×10^9 個引數Imagen|2022|7.8×10^9 個引數GPT-2 Medium (FlashAttention)|2022|3.6×10^8 個引數Tranception|2022|7.0×10^8 個引數CogVideo|2022|9.4×10^9 個引數DITTO|2022|7.5×10^8 個引數CoCa|2022|2.1×10^9 個引數Parti|2022|2.0×10^10 個引數ProGen2-xlarge|2022|6.4×10^9 個引數Minerva (540B)|2022|5.4×10^11 個引數CodeT5-large|2022|7.7×10^8 個引數NLLB|2022|5.5×10^10 個引數BLOOM-176B|2022|1.8×10^11 個引數ESM2-15B|2022|1.5×10^10 個引數OmegaPLM|2022|6.7×10^8 個引數AlexaTM 20B|2022|2.0×10^10 個引數GLM-130B|2022|1.3×10^11 個引數BlenderBot 3|2022|1.8×10^11 個引數BEIT-3|2022|1.9×10^9 個引數PaLI|2022|1.7×10^10 個引數Whisper|2022|1.6×10^9 個引數DiffDock|2022|2.0×10^7 個引數Phenaki|2022|1.8×10^9 個引數GenSLM|2022|2.5×10^10 個引數Flan-PaLM 540B|2022|5.4×10^11 個引數LMSI-Palm|2022|5.4×10^11 個引數U-PaLM (540B)|2022|5.4×10^11 個引數Mogrifier RLSTM (WT2)|2022|3.5×10^7 個引數eDiff-I|2022|9.1×10^9 個引數mT0-13B|2022|1.3×10^10 個引數InternImage|2022|1.1×10^9 個引數EVA-01|2022|1.0×10^9 個引數Galactica|2022|1.2×10^11 個引數AR-LDM|2022|1.5×10^9 個引數Fusion in Encoder|2022|3.3×10^8 個引數ALM 1.0|2022|3.4×10^8 個引數Vega v2|2022|6.0×10^9 個引數RT-1|2022|3.5×10^7 個引數CaLM|2022|8.6×10^7 個引數Hybrid H3-2.7B|2022|2.7×10^9 個引數VALL-E|2023|3.5×10^8 個引數DreamerV3|2023|2.0×10^8 個引數Ankh_large|2023|1.9×10^9 個引數Nucleotide Transformer|2023|2.5×10^9 個引數DDPM-IP (CelebA)|2023|3.0×10^8 個引數MusicLM|2023|8.6×10^8 個引數BLIP-2 (Q-Former)|2023|1.5×10^9 個引數ViT-22B|2023|2.2×10^10 個引數BASIC-L + Lion|2023|3.1×10^9 個引數LLaMA-65B|2023|6.5×10^10 個引數DiT-XL/2|2023|6.8×10^8 個引數AudioGen|2023|1.0×10^9 個引數PaLM-E|2023|5.6×10^11 個引數Falcon-40B|2023|4.0×10^10 個引數GPT-4 (Mar 2023)|2023|1.8×10^12 個引數LEP-AD|2023|3.0×10^9 個引數PanGu-Σ|2023|1.1×10^12 個引數SigLIP 400M|2023|4.0×10^8 個引數BloombergGPT|2023|5.1×10^10 個引數VideoMAE V2|2023|1.0×10^9 個引數Segment Anything Model|2023|6.4×10^8 個引數Incoder-6.7B|2023|6.7×10^9 個引數DINOv2|2023|1.1×10^9 個引數LLaVA|2023|1.3×10^10 個引數ImageBind|2023|9.3×10^8 個引數PaLM 2|2023|3.4×10^11 個引數StarCoder|2023|1.6×10^10 個引數InstructBLIP|2023|1.3×10^10 個引數CoEdiT-xxl|2023|1.1×10^10 個引數Med-PaLM 2|2023|3.4×10^11 個引數CodeT5+|2023|1.6×10^10 個引數ONE-PEACE|2023|4.0×10^9 個引數Goat-7B|2023|7.0×10^9 個引數DPO on Pythia-2.8B|2023|2.8×10^9 個引數PaLI-X|2023|5.5×10^10 個引數MusicGen|2023|3.4×10^9 個引數GPT-3.5 Turbo|2023|2.0×10^10 個引數GPT-4 (Jun 2023)|2023|1.8×10^12 個引數HyenaDNA|2023|6.6×10^6 個引數Stable Diffusion XL (SDXL)|2023|3.4×10^9 個引數InternLM|2023|1.0×10^11 個引數Pangu-Weather|2023|2.6×10^8 個引數xTrimoPGLM -100B|2023|1.0×10^11 個引數GPT3-2.7B (FlashAttention-2)|2023|2.7×10^9 個引數Llama 2-70B|2023|7.0×10^10 個引數Llama 2-7B|2023|7.0×10^9 個引數AudioLM|2023|1.5×10^9 個引數RT-2|2023|5.5×10^10 個引數Qwen-VL|2023|9.6×10^9 個引數Jais|2023|1.3×10^10 個引數Swift|2023|5.7×10^4 個引數Falcon-180B|2023|1.8×10^11 個引數Robot Parkour|2023|5.0×10^5 個引數AlphaMissense|2023|9.3×10^7 個引數Amazon Titan|2023|2.0×10^11 個引數GPT-3.5 Turbo Instruct|2023|2.0×10^10 個引數FinGPT-13B|2023|1.3×10^10 個引數Ferret (13B)|2023|1.3×10^10 個引數RT-2-X|2023|5.5×10^10 個引數PaLI-3|2023|5.0×10^9 個引數CODEFUSION (Python)|2023|7.5×10^7 個引數ChatGLM3-6B|2023|6.0×10^9 個引數DiT-XL/2 + CADS|2023|6.8×10^8 個引數Skywork-13B|2023|1.3×10^10 個引數BLUUMI|2023|1.8×10^11 個引數Grok-1|2023|3.1×10^11 個引數LLaVA 1.5|2023|1.3×10^10 個引數Yi-34B|2023|3.4×10^10 個引數CogVLM-17B|2023|1.7×10^10 個引數RoFormer|2023|1.1×10^8 個引數mPLUG-Owl2|2023|7.1×10^9 個引數Nemotron-3-8B|2023|8.0×10^9 個引數Qwen-Audio-Chat|2023|8.5×10^9 個引數SPHINX (Llama 2 13B)|2023|2.0×10^10 個引數Volcano 13B|2023|1.3×10^10 個引數GNoME for crystal discovery|2023|1.6×10^7 個引數PPLX-70B-Online|2023|7.0×10^10 個引數Qwen-72B|2023|7.2×10^10 個引數Mamba-24M (SC09)|2023|2.3×10^7 個引數Llama Guard|2023|7.0×10^9 個引數SeamlessM4T|2023|2.3×10^9 個引數Mixtral 8x7B|2023|4.7×10^10 個引數W.A.L.T|2023|4.7×10^9 個引數CogAgent|2023|1.8×10^10 個引數FunSearch|2023|1.5×10^10 個引數VILA-13B|2023|1.3×10^10 個引數Gemini Nano-1|2023|1.8×10^9 個引數Gemini Nano-2|2023|3.3×10^9 個引數nekomata-14b|2023|1.4×10^10 個引數GQA-8-XXL|2023|1.1×10^10 個引數CoRe|2023|1.2×10^10 個引數Palmyra X 003|2024|7.2×10^10 個引數AlphaGeometry|2024|1.5×10^8 個引數Qwen-VL-Max|2024|7.0×10^9 個引數Qwen1.5-72B|2024|7.2×10^10 個引數Aya|2024|1.3×10^10 個引數MegaScale (Production)|2024|5.3×10^11 個引數Stable Diffusion 3|2024|8.0×10^9 個引數Aramco Metabrain AI|2024|2.5×10^11 個引數MM1-30B|2024|3.0×10^10 個引數DBRX|2024|1.3×10^11 個引數ReALM|2024|3.0×10^9 個引數Reka Core|2024|6.7×10^10 個引數Llama 3-70B|2024|7.0×10^10 個引數VILA1.5-13B|2024|1.4×10^10 個引數Yi-Large|2024|1.0×10^11 個引數Octo-Base|2024|9.3×10^7 個引數ALLaM adapted 70B|2024|7.0×10^10 個引數Qwen2-72B|2024|7.3×10^10 個引數Nemotron-4 340B|2024|3.4×10^11 個引數OpenVLA|2024|7.2×10^9 個引數DeepSeek-Coder-V2 236B|2024|2.4×10^11 個引數Cambrian-1-34B|2024|3.4×10^10 個引數ESM3 (98B)|2024|9.9×10^10 個引數SenseChat 5.5|2024|6.0×10^11 個引數Mathstral|2024|7.0×10^9 個引數Llama 3.1-405B|2024|4.1×10^11 個引數Mistral Large 2|2024|1.2×10^11 個引數AFM-on-device|2024|2.7×10^9 個引數LLaVA-OV-72B|2024|7.2×10^10 個引數Table Tennis Agent|2024|1.9×10^5 個引數Jamba 1.5-Large|2024|4.0×10^11 個引數DeepSeek-V2.5|2024|2.4×10^11 個引數Qwen2.5-32B|2024|3.3×10^10 個引數Oryx 34B|2024|3.4×10^10 個引數Qwen2.5 Instruct (72B)|2024|7.3×10^10 個引數Qwen2.5-72B|2024|7.3×10^10 個引數Telechat2-115B|2024|1.2×10^11 個引數Llama 3.2 11B|2024|1.1×10^10 個引數Movie Gen Video|2024|3.0×10^10 個引數GR-2|2024|2.3×10^8 個引數Palmyra X 004|2024|1.5×10^11 個引數RDT-1B|2024|1.2×10^9 個引數NVLM-D 72B|2024|7.2×10^10 個引數NVLM-H 72B|2024|7.2×10^10 個引數NVLM-X 72B|2024|7.2×10^10 個引數Doubao-pro|2024|5.0×10^11 個引數Hunyuan-Large|2024|3.9×10^11 個引數Pixtral Large|2024|1.2×10^11 個引數Fugatto 1|2024|2.5×10^9 個引數未知模型|2024|2.0×10^9 個引數Llama 3.3 70B|2024|7.0×10^10 個引數NVILA 15B|2024|1.5×10^10 個引數EXAONE 3.5 32B|2024|3.2×10^10 個引數Apollo 7B|2024|7.0×10^9 個引數DeepSeek-V3|2024|6.7×10^11 個引數STORM-B/8|2025|1.0×10^8 個引數INTELLECT-MATH|2025|7.0×10^9 個引數DeepSeek-R1|2025|6.7×10^11 個引數Eagle 2|2025|8.9×10^9 個引數Eurus-2-7B-PRIME|2025|7.0×10^9 個引數Grok 3|2025|3.0×10^12 個引數QwQ-32B|2025|3.3×10^10 個引數Hunyuan-TurboS|2025|5.6×10^11 個引數ERNIE-4.5-VL-424B-A47B (文心大模型4.5)|2025|4.2×10^11 個引數EXAONE Deep 32B|2025|3.2×10^10 個引數DeepSeek-V3 (Mar 2025)|2025|6.7×10^11 個引數Diffusion Renderer|2025|1.1×10^9 個引數Llama 4 Behemoth (preview)|2025|2.0×10^12 個引數Llama 4 Maverick|2025|4.0×10^11 個引數Llama 4 Scout|2025|1.1×10^11 個引數Pangu Ultra|2025|1.3×10^11 個引數Qwen3-235B-A22B|2025|2.4×10^11 個引數DeepSeek-R1 (May 2025)|2025|6.7×10^11 個引數Qwen3 Embedding|2025|8.0×10^9 個引數FGN|2025|7.2×10^8 個引數Seed-1.6-Thinking|2025|2.3×10^11 個引數EXAONE Path 2.0|2025|1.8×10^8 個引數Grok 4|2025|3.0×10^12 個引數Kimi K2|2025|1.0×10^12 個引數EXAONE 4.0 (32B)|2025|3.2×10^10 個引數Qwen3-Coder-480B-A35B|2025|4.8×10^11 個引數Qwen3-235B-A22B (Jul 2025)|2025|2.4×10^11 個引數Qwen3-235B-A22B-Thinking (Jul 2025)|2025|2.4×10^11 個引數MindLink-72B|2025|7.2×10^10 個引數GLM-4.5|2025|3.6×10^11 個引數Hierarchical Reasoning Model (HPM)|2025|2.7×10^7 個引數Qwen Image|2025|2.7×10^10 個引數gpt-oss-120b|2025|1.2×10^11 個引數gpt-oss-20b|2025|2.1×10^10 個引數LongCat-Flash|2025|5.6×10^11 個引數Qwen3-Max|2025|1.0×10^12 個引數AgentFounder-30B|2025|3.0×10^10 個引數Qwen3-Omni-30B-A3B|2025|3.5×10^10 個引數GLM-4.6|2025|3.6×10^11 個引數Ling-1T|2025|1.0×10^12 個引數MiniMax-M2|2025|2.3×10^11 個引數Tongyi DeepResearch|2025|3.1×10^10 個引數Kaiju Large|2025|1.1×10^11 個引數Kaiju Medium|2025|3.4×10^10 個引數Kaiju Small|2025|1.3×10^10 個引數Kimi K2 Thinking|2025|1.0×10^12 個引數Olmo 3|2025|3.2×10^10 個引數P1-235B-A22B|2025|2.4×10^11 個引數π0.6 (pi-0.6)|2025|5.3×10^9 個引數DeepSeekMath-V2|2025|6.9×10^11 個引數Nemotron 3-Nano-30B-A3B|2025|3.2×10^10 個引數GLM-4.7|2025|3.6×10^11 個引數MiniMax-M2.1|2025|2.3×10^11 個引數A.X K1|2025|5.2×10^11 個引數HyperCLOVA X SEED 32B Think|2025|3.2×10^10 個引數VAETKI|2025|1.0×10^11 個引數K-EXAONE|2026|2.4×10^11 個引數Solar Open 100B|2026|1.0×10^11 個引數Kimi K2.5|2026|1.0×10^12 個引數Qwen3-Coder-Next|2026|8.0×10^10 個引數GLM-5|2026|7.4×10^11 個引數Qwen3.5 397B-A17B|2026|4.0×10^11 個引數Grok 4.20|2026|5.0×10^11 個引數Qwen3.5-122B-A10B|2026|1.2×10^11 個引數Nemotron 3 Super|2026|1.2×10^11 個引數Composer 2|2026|1.0×10^12 個引數MiMo-V2-Pro|2026|1.0×10^12 個引數GLM-5.1|2026|7.5×10^11 個引數MolmoAct 2|2026|5.5×10^9 個引數EXAONE 4.5|2026|3.3×10^10 個引數Kimi K2.6|2026|1.0×10^12 個引數DeepSeek-V4-Flash|2026|2.8×10^11 個引數DeepSeek-V4-Pro|2026|1.6×10^12 個引數MiMo-V2.5-Pro|2026|1.0×10^12 個引數Tencent Hy3 preview|2026|3.0×10^11 個引數TML-Interaction-Small|2026|2.8×10^11 個引數Composer 2.5|2026|1.0×10^12 個引數Nemotron 3 Ultra|2026|5.5×10^11 個引數Solar Open2 250B|2026|2.5×10^11 個引數Tencent Hy3|2026|3.0×10^11 個引數Inkling|2026|9.8×10^11 個引數Kimi K3|2026|2.8×10^12 個引數Qwen 3.8 Max|2026|2.4×10^12 個引數A.X K2|2026|6.9×10^11 個引數K-EXAONE 2.0|2026|7.5×10^11 個引數Motif-3|2026|3.1×10^11 個引數DeepSeek-V4-Pro-0813|2026|1.6×10^12 個引數GLM-5.3|2026|7.4×10^11 個引數其他
每個點為一個知名 AI 模型,縱軸=可訓練引數量(對數軸),按開發方型別著色;共 718 個模型。

訓練資料量趨勢

訓練資料集的樣本 / token 規模——喂進模型的樣本 / token 規模。資料量與算力、引數量一同增長,是規模定律的第三根支柱。

110³10⁶10⁹10¹²10¹⁵1960198020002020釋出年份訓練資料量(個樣本,對數軸)Theseus|1950|4.0×10^1 個樣本Self Organizing System|1955|2.0×10^0 個樣本Perceptron Mark I|1957|1.0×10^2 個樣本Pattern recognition and reading by machine|1959|1.8×10^2 個樣本Perceptron (1960)|1960|1.0×10^2 個樣本ADALINE|1960|1.0×10^2 個樣本Linear Decision Functions|1962|5.0×10^2 個樣本MADALINE I|1962|2.6×10^2 個樣本LTE speaker verification system|1966|4.2×10^2 個樣本GLEE|1968|6.0×10^3 個樣本Piecewise linear model|1973|3.1×10^2 個樣本Cognitron|1975|5.0×10^0 個樣本Neocognitron|1980|5.0×10^0 個樣本Kohonen network|1981|4.0×10^3 個樣本ASE+ACE|1983|5.0×10^5 個樣本Hierarchical Cognitron|1984|5.0×10^0 個樣本Error Propagation|1986|6.4×10^1 個樣本Distributed representation NN|1986|1.0×10^2 個樣本MLP with back-propagation|1986|1.0×10^2 個樣本NetTalk (dictionary)|1987|5.0×10^3 個樣本NetTalk (transcription)|1987|5.1×10^3 個樣本Translation-invariant MLP|1987|1.6×10^2 個樣本MLN-ASR|1988|1.3×10^4 個樣本MLP baggage detector|1989|2.0×10^4 個樣本Q-learning|1989|2.0×10^5 個樣本Handwritten digit recognition network|1989|9.8×10^3 個樣本Speaker-independent vowel classification|1989|4.1×10^3 個樣本Zip CNN|1989|7.3×10^3 個樣本NETtalk reimplementation|1990|7.2×10^3 個樣本Bankruptcy-NN|1990|7.4×10^1 個樣本ISR network|1990|6.0×10^5 個樣本SexNet classification|1990|8.0×10^1 個樣本SexNet compression|1990|8.1×10^4 個樣本RAAM|1990|2.9×10^1 個樣本Weight Decay|1991|2.5×10^4 個樣本TD-Gammon|1992|6.3×10^6 個樣本Golem|1992|1.6×10^3 個樣本Cancer drug mechanism prediction|1992|1.4×10^2 個樣本Boosting|1992|2.9×10^4 個樣本IBM-5|1993|2.9×10^7 個樣本Siamese-TDNN|1993|7.7×10^3 個樣本ANN Eye Tracker|1993|4.0×10^3 個樣本Ceramic-MLP|1994|8.0×10^1 個樣本JPMAX|1994|1.5×10^3 個樣本Mixture of linear models|1994|1.8×10^6 個樣本NeuroChess|1994|9.6×10^6 個樣本Predictive Coding NN|1994|6.0×10^5 個樣本Support Vector Machines|1995|6.0×10^4 個樣本LISSOM|1995|2.0×10^3 個樣本MUSIC perceptron|1996|8.1×10^4 個樣本System 11|1996|2.4×10^4 個樣本AdaBoost.M2 Digit Recognition|1996|9.7×10^3 個樣本SOM-CNN|1997|1.3×10^5 個樣本Bidirectional RNN|1997|1.4×10^5 個樣本LSTM|1997|8.5×10^5 個樣本LeNet-5|1998|6.0×10^4 個樣本LSTM with forget gates|1999|1.4×10^8 個樣本RECONTRA-categorized|1999|4.0×10^4 個樣本RECONTRA-uncategorized|1999|5.8×10^4 個樣本IBM Model 4|1999|8.0×10^5 個樣本Neural LM|2000|3.2×10^7 個樣本PoE MNIST|2000|5.4×10^4 個樣本Gradient Boosting Machine|2001|5.0×10^3 個樣本Decision tree (classification)|2001|7.5×10^5 個樣本Thumbs Up?|2002|1.4×10^3 個樣本NPLM (AP News)|2003|1.4×10^7 個樣本NPLM (Brown)|2003|1.4×10^7 個樣本Invariant CNN|2004|2.4×10^4 個樣本LMICA|2004|1.0×10^5 個樣本Hierarchical LM|2005|9.0×10^5 個樣本Histograms of Oriented Gradients|2005|1.5×10^4 個樣本RankNet|2005|3.5×10^6 個樣本TFE SVM|2006|6.0×10^5 個樣本SVM-CNN|2006|5.8×10^5 個樣本Spatial Pyramid Matching|2006|3.0×10^3 個樣本Deep Belief Nets|2006|4.7×10^7 個樣本Dimensionality Reduction|2006|4.7×10^7 個樣本Greedy layer-wise DNN training|2006|1.1×10^8 個樣本Local Binary Patterns for facial recognition|2006|7.4×10^2 個樣本KN-LM|2007|3.1×10^10 個樣本SB-LM|2007|1.8×10^12 個樣本BLSTM for handwriting (1)|2007|4.1×10^5 個樣本Enhanced Neighborhood-Based Filtering|2007|1.0×10^8 個樣本BLSTM for handwriting (2)|2007|3.3×10^6 個樣本Deep Multitask NLP Network|2008|6.3×10^8 個樣本Denoising Autoencoders|2008|7.8×10^6 個樣本HLBL|2008|1.4×10^7 個樣本GNN|2008|2.1×10^2 個樣本RBM Image Classifier|2009|6.1×10^9 個樣本GPU DBNs|2009|1.2×10^11 個樣本MatrixFac for Recommenders|2009|1.0×10^8 個樣本Two Stage Feature Extraction (MNIST)|2009|5.0×10^4 個樣本LCNP LabelMe|2009|4.0×10^4 個樣本LCNP MNIST|2009|5.0×10^4 個樣本LCNP NORB|2009|2.4×10^4 個樣本Stacked Denoising Autoencoders|2010|3.4×10^8 個樣本Feedforward NN|2010|9.0×10^4 個樣本ReLU (LFW)|2010|2.3×10^5 個樣本ReLU (NORB)|2010|2.9×10^5 個樣本iCCCP|2010|1.0×10^4 個樣本Pooling CNN (Caltech 101)|2010|3.1×10^3 個樣本Pooling CNN (NORB)|2010|2.4×10^4 個樣本RNN LM|2010|6.4×10^6 個樣本Deep rectifier networks|2011|8.2×10^7 個樣本Deep Autoencoders|2011|4.9×10^9 個樣本Vector Space Model|2011|5.7×10^6 個樣本Recursive Neural Network|2011|5.7×10^5 個樣本High Performance CNN (NORB)|2011|5.0×10^4 個樣本CNN Committee (MNIST)|2011|4.2×10^5 個樣本CNN Committee (NIST)|2011|3.4×10^6 個樣本Adaptive Subgrad|2011|8.0×10^5 個樣本CNN committee (traffic sign)|2011|5.3×10^4 個樣本NLP from scratch|2011|8.5×10^8 個樣本Dropout (CIFAR)|2012|6.0×10^4 個樣本Dropout (ImageNet)|2012|2.6×10^6 個樣本Dropout (MNIST)|2012|6.0×10^4 個樣本Unsupervised High-level Feature Learner|2012|1.2×10^12 個樣本Context-dependent RNN|2012|3.7×10^7 個樣本LSTM LM|2012|2.7×10^7 個樣本AlexNet|2012|2.5×10^9 個樣本Bayesian automated hyperparameter tuning|2012|5.0×10^4 個樣本DNN EM segmentation|2012|3.0×10^6 個樣本DistBelief Speech|2012|1.1×10^9 個樣本DistBelief Vision|2012|1.6×10^7 個樣本RNN+LDA+KN5+cache|2012|9.3×10^5 個樣本DistBelief NNLM|2013|6.0×10^9 個樣本Multilingual DNN|2013|3.1×10^9 個樣本Hierarchical Scene Labeling (Stanford Background)|2013|7.1×10^7 個樣本RCTM|2013|4.5×10^6 個樣本RNTN|2013|1.6×10^5 個樣本Word2Vec (large)|2013|3.3×10^11 個樣本Word2Vec (small)|2013|1.0×10^10 個樣本Visualizing CNNs|2013|7.7×10^6 個樣本DeViSE|2013|5.4×10^9 個樣本TransE|2013|1.8×10^7 個樣本RNN for 1B words|2013|1.0×10^9 個樣本DQN|2013|1.6×10^8 個樣本Network in Network|2013|6.3×10^5 個樣本Image generation|2013|4.7×10^7 個樣本GloVe (32B)|2014|3.2×10^8 個樣本GloVe (6B)|2014|6.6×10^7 個樣本HyperNEAT|2014|7.5×10^8 個樣本Paragraph Vector|2014|1.6×10^7 個樣本AdaRNN|2014|6.3×10^3 個樣本Dropout: SVHN|2014|6.0×10^5 個樣本GANs|2014|1.2×10^5 個樣本Two-stream ConvNets for action recognition|2014|1.3×10^6 個樣本SPPNet|2014|1.3×10^6 個樣本DeepFace|2014|4.4×10^6 個樣本Fragment embedding|2014|1.5×10^7 個樣本Multiresolution CNN|2014|5.0×10^7 個樣本ACF-WIDER|2014|1.4×10^5 個樣本NPD|2014|4.4×10^5 個樣本RNNsearch-50*|2014|2.3×10^8 個樣本VGG16|2014|1.3×10^6 個樣本VGG19|2014|1.3×10^6 個樣本Seq2Seq LSTM|2014|8.7×10^8 個樣本SPN-4+KN5|2014|9.3×10^5 個樣本Deeply-supervised nets|2014|6.0×10^5 個樣本GoogLeNet / InceptionV1|2014|5.7×10^11 個樣本Spatially-Sparse CNN|2014|9.0×10^5 個樣本LRCN|2014|4.0×10^5 個樣本SC-NLM|2014|5.0×10^6 個樣本Cascaded LNet-ANet|2014|9.3×10^6 個樣本TA-CNN|2014|4.5×10^4 個樣本SNM-skip|2014|8.0×10^8 個樣本Fractional Max-Pooling|2014|9.0×10^5 個樣本ADAM (CIFAR-10)|2014|5.0×10^4 個樣本VGG-Face|2015|2.6×10^6 個樣本MSRA (C, PReLU)|2015|1.3×10^6 個樣本DQN-2015|2015|1.2×10^7 個樣本genCNN + dyn eval|2015|9.3×10^5 個樣本TC-DNN-BLSTM-DNN|2015|2.9×10^7 個樣本Fast R-CNN|2015|2.6×10^7 個樣本U-Net|2015|7.9×10^6 個樣本Faster R-CNN|2015|1.0×10^8 個樣本CFSS|2015|1.4×10^5 個樣本BatchNorm|2015|1.2×10^10 個樣本Deep CNN + COTS|2015|4.9×10^5 個樣本DCNN|2015|4.9×10^5 個樣本BPE|2015|5.0×10^7 個樣本AlphaGo Fan|2015|1.3×10^10 個樣本SAF R-CNN|2015|3.5×10^5 個樣本3DDFA|2015|2.9×10^5 個樣本Inception v3|2015|1.2×10^6 個樣本SSD|2015|2.3×10^6 個樣本ResNet-101 (ImageNet)|2015|1.3×10^6 個樣本ResNet-110 (CIFAR-10)|2015|5.0×10^4 個樣本ResNet-152 (ImageNet)|2015|1.3×10^6 個樣本Advantage Learning|2015|1.0×10^8 個樣本Variational (untied weights, MC) LSTM (Large)|2015|9.3×10^5 個樣本AlphaGo Lee|2016|3.0×10^8 個樣本A3C FF hs|2016|2.0×10^8 個樣本Inception-ResNet-V2|2016|1.3×10^6 個樣本Inceptionv4|2016|1.3×10^6 個樣本SqueezeNet|2016|1.3×10^6 個樣本Named Entity Recognition model|2016|2.1×10^5 個樣本Template Adaptation|2016|7.8×10^3 個樣本Gated HORNN (3rd order)|2016|2.2×10^7 個樣本LRR-4X|2016|1.5×10^8 個樣本PixelCNN|2016|1.6×10^10 個樣本R-FCN|2016|1.1×10^7 個樣本CCL|2016|2.0×10^4 個樣本SimpleNet|2016|1.3×10^6 個樣本LF-MMI|2016|7.2×10^5 個樣本MS-ensemble-speech-recognition|2016|1.1×10^10 個樣本WaveNet|2016|1.2×10^10 個樣本ResNet-1001|2016|5.0×10^4 個樣本ResNet-200|2016|1.3×10^6 個樣本Wide Residual Network|2016|1.3×10^6 個樣本GNMT|2016|7.2×10^8 個樣本Pointer Sentinel-LSTM (medium)|2016|9.3×10^5 個樣本GAWWN|2016|2.4×10^5 個樣本Xception|2016|3.5×10^8 個樣本SPIDER2|2016|1.4×10^7 個樣本BIDAF|2016|8.8×10^5 個樣本NAS with base 8 and shared embeddings|2016|9.3×10^5 個樣本NASv3 (CIFAR-10)|2016|4.5×10^4 個樣本VD-LSTM+REAL Large|2016|9.3×10^5 個樣本DLDL (PASCAL)|2016|2.3×10^4 個樣本DTN (Domain Transfer Network)|2016|2.0×10^6 個樣本DAC-CSR|2016|2.0×10^4 個樣本ResNeXt-101 (64×4d)|2016|1.3×10^6 個樣本ResNeXt-50|2016|1.3×10^6 個樣本PolyNet|2016|1.3×10^6 個樣本Image-to-image cGAN|2016|2.4×10^6 個樣本PointNet|2016|9.8×10^3 個樣本3DMM-CNN|2016|5.0×10^5 個樣本HR-ResNet101|2016|8.2×10^6 個樣本EnhanceNet|2016|9.8×10^9 個樣本YOLOv2|2016|1.3×10^6 個樣本DeepStack|2017|2.5×10^10 個樣本OR-WideResNet|2017|5.0×10^4 個樣本MoE-Multi|2017|8.7×10^10 個樣本DnCNN|2017|2.6×10^9 個樣本Prototypical networks|2017|3.8×10^4 個樣本Mask R-CNN|2017|4.6×10^10 個樣本MobileNet|2017|1.3×10^6 個樣本DeepLab (2017)|2017|2.6×10^7 個樣本Mnemonic Reader|2017|2.2×10^5 個樣本SRGAN|2017|7.0×10^5 個樣本Inflated 3D ConvNet|2017|2.4×10^5 個樣本PointNet++|2017|6.0×10^4 個樣本Reading Twice for NLU|2017|2.0×10^5 個樣本Transformer (2017)|2017|8.3×10^8 個樣本HRA|2017|1.5×10^8 個樣本DeepLabV3|2017|8.4×10^9 個樣本NoisyNet-Dueling|2017|3.2×10^8 個樣本ShuffleNet v1|2017|1.3×10^6 個樣本JFT|2017|5.5×10^12 個樣本AWD-LSTM|2017|2.0×10^6 個樣本NASNet-A|2017|1.3×10^6 個樣本ConvS2S (ensemble of 8 models)|2017|1.2×10^9 個樣本GSM|2017|2.2×10^5 個樣本AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2)|2017|2.0×10^6 個樣本RetinaNet-R101|2017|1.2×10^5 個樣本RetinaNet-R50|2017|1.2×10^10 個樣本EI-REHN-1000D|2017|9.3×10^5 個樣本NeuMF (Pinterest)|2017|1.5×10^6 個樣本GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2)|2017|2.0×10^6 個樣本PyramidNet|2017|1.3×10^6 個樣本SENet (ImageNet)|2017|1.3×10^6 個樣本ISS|2017|9.3×10^5 個樣本LSTM + dynamic eval|2017|9.0×10^7 個樣本AWD-LSTM+WT+Cache+IOG (WT2)|2017|2.0×10^6 個樣本AlphaGo Zero|2017|6.4×10^9 個樣本PhraseCond|2017|1.6×10^5 個樣本S-Norm|2017|1.1×10^6 個樣本DCN+|2017|2.2×10^5 個樣本Fraternal dropout + AWD-LSTM 3-layer (WT2)|2017|2.0×10^6 個樣本VQ-VAE|2017|6.3×10^10 個樣本AWD-LSTM-MoS + dynamic evaluation (WT2, 2017)|2017|2.0×10^6 個樣本TriNet|2017|5.1×10^5 個樣本DL scaling LM|2017|4.0×10^8 個樣本DL scaling speech|2017|2.2×10^9 個樣本AlphaZero|2017|3.5×10^9 個樣本ELMo|2018|2.0×10^9 個樣本QRNN|2018|1.0×10^8 個樣本T-DMCA|2018|1.4×10^10 個樣本DeepLabV3+|2018|8.7×10^9 個樣本IMPALA|2018|1.1×10^10 個樣本TCN (P-MNIST)|2018|6.0×10^4 個樣本4 layer QRNN (h=2500)|2018|1.0×10^8 個樣本YOLOv3|2018|5.4×10^6 個樣本Dropout-LSTM+Noise(Bernoulli) (WT2)|2018|2.0×10^6 個樣本ResNeXt-101 32x48d|2018|9.4×10^8 個樣本aLSTM(depth-2)+RecurrentPolicy (WT2)|2018|2.0×10^6 個樣本GPT-1|2018|1.3×10^9 個樣本Relational Memory Core|2018|4.0×10^9 個樣本MobileNetV2|2018|1.3×10^6 個樣本FTW (For The Win)|2018|2.0×10^9 個樣本Big-Little Net|2018|1.3×10^6 個樣本Big-Little Net (speech)|2018|7.2×10^8 個樣本AWD-LSTM-MoS+PDR + dynamic evaluation (WT2)|2018|2.0×10^6 個樣本Big Transformer for Back-Translation|2018|4.5×10^9 個樣本(ensemble): AWD-LSTM-DOC (fin) × 5 (WT2)|2018|2.0×10^6 個樣本AWD-LSTM-MoS + dynamic evaluation (WT2, 2018)|2018|2.0×10^6 個樣本Transformer + Simple Recurrent Unit|2018|1.1×10^8 個樣本LSTM+NeuralCache|2018|2.0×10^6 個樣本Transformer (Adaptive Input Embeddings) WT103|2018|1.0×10^8 個樣本BERT-Large|2018|2.7×10^9 個樣本TrellisNet|2018|1.0×10^8 個樣本MemoReader|2018|1.1×10^6 個樣本Mesh-TensorFlow Transformer 2.9B (translation)|2018|1.6×10^9 個樣本Mesh-TensorFlow Transformer 4.9B (language)|2018|5.0×10^9 個樣本Fine-tuned-AWD-LSTM-DOC (fin)|2018|1.0×10^6 個樣本GPipe (Transformer)|2018|1.5×10^12 個樣本Multi-cell LSTM|2018|9.3×10^5 個樣本SPN (ImageNet 128)|2018|2.5×10^11 個樣本StyleGAN|2018|5.0×10^7 個樣本Transformer ELMo|2019|2.0×10^9 個樣本Transformer-XL (257M)|2019|1.0×10^8 個樣本Hanabi 4 player|2019|2.0×10^10 個樣本MT-DNN|2019|1.0×10^6 個樣本GPT-2 (1.5B)|2019|1.1×10^10 個樣本KataGo|2019|2.4×10^8 個樣本SciBERT|2019|3.2×10^9 個樣本True-Regularization+Finetune+Dynamic-Eval|2019|9.3×10^5 個樣本WeNet (Penn Treebank)|2019|9.3×10^5 個樣本Transformer-XL + RMS dynamic eval|2019|1.0×10^8 個樣本BERT-Large-CAS (PTB+WT2+WT103)|2019|1.3×10^9 個樣本Neuro-Symbolic Concept Learner|2019|1.0×10^5 個樣本ResNeXt-101 Billion-scale|2019|9.0×10^7 個樣本AWD-LSTM-DRILL + dynamic evaluation† (WT2)|2019|2.0×10^6 個樣本EfficientNet-L2|2019|1.3×10^6 個樣本DLRM-2020|2019|3.9×10^7 個樣本XLNet|2019|3.3×10^10 個樣本Transformer-XL Large + Phrase Induction|2019|1.0×10^8 個樣本AWD-LSTM + MoS + Partial Shuffled|2019|2.0×10^6 個樣本Char-CNN-BiLSTM|2019|9.3×10^5 個樣本FixRes ResNeXt-101 WSL|2019|9.4×10^8 個樣本LaNet-L (CIFAR-10)|2019|6.0×10^4 個樣本BigBiGAN|2019|2.6×10^6 個樣本RoBERTa Large|2019|4.3×10^10 個樣本Mogrifier (d2, MoS2, MC) + dynamic eval|2019|2.0×10^6 個樣本UDSMProt|2019|1.5×10^8 個樣本Megatron-BERT|2019|7.0×10^9 個樣本Megatron-LM (1.2B)|2019|1.6×10^11 個樣本Megatron-LM (8.3B)|2019|4.6×10^10 個樣本ALBERT|2019|3.3×10^9 個樣本Adaptive Inputs + LayerDrop|2019|1.0×10^8 個樣本AlphaX-1|2019|6.1×10^7 個樣本DistilBERT|2019|5.0×10^8 個樣本T5-11B|2019|3.4×10^10 個樣本T5-3B|2019|5.1×10^9 個樣本BART-large|2019|4.3×10^10 個樣本Base LM + kNN LM + Continuous Cache|2019|1.0×10^8 個樣本XLM-RoBERTa|2019|1.7×10^11 個樣本CamemBERT|2019|2.9×10^10 個樣本Noisy Student (L2)|2019|8.1×10^7 個樣本Sandwich Transformer|2019|7.0×10^8 個樣本MoCo|2019|9.4×10^8 個樣本MuZero|2019|1.2×10^10 個樣本Transformer - LibriVox + Decoding/Rescoring|2019|9.8×10^8 個樣本Photo-Geometric Autoencoder|2019|8.2×10^8 個樣本Transformer-XL DeFINE (141M)|2019|1.0×10^8 個樣本StarGAN v2|2019|4.0×10^5 個樣本StyleGAN2|2019|1.1×10^8 個樣本MMLSTM (PTB)|2019|9.3×10^5 個樣本MMLSTM (WT-2)|2019|2.0×10^6 個樣本OpenAI Five|2019|4.5×10^11 個樣本OpenAI Five Rerun|2019|5.3×10^10 個樣本DD-PPO|2019|2.5×10^9 個樣本Big Transfer (BiT-L)|2019|3.0×10^8 個樣本AlphaFold|2020|6.6×10^9 個樣本Meena|2020|5.3×10^10 個樣本Perceiver IO (optical flow)|2020|1.5×10^11 個樣本TaLK Convolution|2020|1.0×10^8 個樣本Theseus 6/768|2020|3.9×10^5 個樣本ALBERT-xxlarge|2020|3.3×10^9 個樣本FFN SwiGLU|2020|5.1×10^10 個樣本SimCLR|2020|1.1×10^10 個樣本Turing-NLG|2020|4.6×10^10 個樣本Feedback Transformer|2020|1.0×10^8 個樣本TCAN (WT2)|2020|2.0×10^6 個樣本Routing Transformer (WT-103)|2020|1.0×10^8 個樣本TransformerXL + spectrum control|2020|1.0×10^8 個樣本Tensor-Transformer(1core)+PN (WT103)|2020|1.0×10^8 個樣本ELECTRA|2020|3.3×10^10 個樣本MetNet|2020|7.1×10^9 個樣本Go-explore|2020|4.0×10^10 個樣本Once for All|2020|1.3×10^6 個樣本ContextNet|2020|3.5×10^8 個樣本Retrieval-Augmented Generator|2020|3.1×10^6 個樣本DETR|2020|8.3×10^5 個樣本GPT-3 175B (davinci)|2020|2.4×10^11 個樣本GShard (dense)|2020|3.5×10^11 個樣本DeLighT|2020|1.0×10^8 個樣本ERNIE-GEN (large)|2020|1.2×10^11 個樣本ProBERTa|2020|5.8×10^7 個樣本LUKE|2020|4.7×10^9 個樣本German ELECTRA Large|2020|3.6×10^10 個樣本mT5-XXL|2020|1.0×10^12 個樣本ViT-Base/32|2020|3.0×10^8 個樣本ViT-Huge/14|2020|3.0×10^8 個樣本wave2vec 2.0 LARGE|2020|4.6×10^9 個樣本KEPLER|2020|3.5×10^9 個樣本AlphaFold 2|2020|5.7×10^9 個樣本CPM-Large|2020|1.7×10^10 個樣本ESM1b|2020|2.8×10^10 個樣本VQGAN + CLIP|2020|2.5×10^11 個樣本CT-MoS (WT2)|2020|2.0×10^6 個樣本DensePhrases|2020|5.8×10^7 個樣本ERNIE-Doc (247M)|2021|1.0×10^8 個樣本CLIP (ResNet-50)|2021|4.0×10^8 個樣本CLIP (ViT L/14@336px)|2021|4.0×10^8 個樣本DALL-E|2021|3.2×10^11 個樣本BigSSL|2021|1.0×10^11 個樣本Switch|2021|8.6×10^10 個樣本DeiT-B|2021|3.8×10^6 個樣本top-down frozen classifier|2021|3.4×10^6 個樣本MSA Transformer|2021|1.4×10^12 個樣本SRU++ Large|2021|1.0×10^8 個樣本Meta Pseudo Labels|2021|1.3×10^8 個樣本Generative BST|2021|5.7×10^10 個樣本M6-T|2021|1.1×10^11 個樣本PLUG|2021|6.0×10^10 個樣本ProtBERT-BFD|2021|5.9×10^10 個樣本ProtT5-XL-U50|2021|2.0×10^10 個樣本ADM|2021|1.3×10^14 個樣本MedBERT|2021|1.5×10^10 個樣本ByT5-XXL|2021|1.1×10^12 個樣本CogView|2021|9.7×10^11 個樣本Transformer local-attention (NesT-B)|2021|1.3×10^6 個樣本ViT-G/14|2021|3.0×10^9 個樣本ALIGN|2021|1.8×10^9 個樣本CoAtNet|2021|8.9×10^13 個樣本DeBERTa|2021|2.1×10^10 個樣本Denoising Diffusion Probabilistic Models (LSUN Bedroom)|2021|6.0×10^11 個樣本EMDR|2021|1.7×10^11 個樣本StyleGAN3-R|2021|5.0×10^7 個樣本StyleGAN3-T|2021|5.0×10^7 個樣本Fold2Seq|2021|4.6×10^4 個樣本Adaptive Input Transformer + RD|2021|1.0×10^8 個樣本Codex|2021|5.3×10^10 個樣本ERNIE 3.0|2021|3.8×10^11 個樣本GOAT|2021|8.0×10^14 個樣本HuBERT|2021|8.6×10^8 個樣本SEER|2021|1.0×10^9 個樣本6-Act Tether|2021|1.3×10^8 個樣本YOLOX-X|2021|2.5×10^6 個樣本Jurassic-1-Jumbo|2021|3.0×10^11 個樣本DNABERT|2021|1.4×10^9 個樣本XLMR-XXL|2021|1.7×10^11 個樣本FLAN 137B|2021|2.5×10^12 個樣本MEB|2021|5.0×10^11 個樣本PermuteFormer|2021|1.0×10^8 個樣本HyperCLOVA 204B|2021|5.6×10^11 個樣本PLATO-XL|2021|1.5×10^11 個樣本AlphaFold-Multimer|2021|5.7×10^7 個樣本Megatron-Turing NLG 530B|2021|2.7×10^11 個樣本Yuan 1.0|2021|1.8×10^11 個樣本base LM+GNN+kNN|2021|1.0×10^8 個樣本Eve|2021|2.4×10^10 個樣本EfficientZero|2021|1.0×10^5 個樣本Projected GAN|2021|3.0×10^6 個樣本S4|2021|1.0×10^8 個樣本ViT-G/14 (LiT)|2021|1.0×10^12 個樣本AliceMind-MMU|2021|1.3×10^7 個樣本BASIC-L|2021|8.9×10^12 個樣本Florence|2021|7.5×10^9 個樣本NÜWA|2021|5.6×10^9 個樣本Gopher (280B)|2021|3.0×10^11 個樣本Student of Games|2021|2.5×10^11 個樣本GLaM|2021|6.0×10^11 個樣本LongT5|2021|5.2×10^11 個樣本Contriever|2021|2.6×10^11 個樣本LDM-1.45B|2021|2.9×10^11 個樣本XGLM-7.5B|2021|5.0×10^11 個樣本ERNIE 3.0 Titan|2021|6.7×10^11 個樣本data2vec (language)|2022|1.3×10^11 個樣本data2vec (speech)|2022|1.8×10^7 個樣本data2vec (vision)|2022|2.5×10^8 個樣本OntoProtein|2022|2.9×10^9 個樣本InstructGPT 175B|2022|1.7×10^7 個樣本AlphaCode|2022|9.7×10^11 個樣本RETRO-7B|2022|4.2×10^11 個樣本GPT-NeoX-20B|2022|3.4×10^11 個樣本LaMDA|2022|2.1×10^12 個樣本ProteinBERT|2022|3.8×10^10 個樣本ST-MoE|2022|1.5×10^12 個樣本PolyCoder|2022|3.9×10^10 個樣本DeepNet|2022|2.7×10^11 個樣本Statement Curriculum Learning|2022|3.7×10^11 個樣本ViT-G (model soup)|2022|1.8×10^9 個樣本Make-A-Scene|2022|2.7×10^11 個樣本Segatron-XL large, M=384 + HCP|2022|1.0×10^8 個樣本Chinchilla|2022|1.4×10^12 個樣本PaLM (540B)|2022|7.8×10^11 個樣本DALL·E 2|2022|1.7×10^11 個樣本Sparse all-MLP|2022|1.0×10^11 個樣本Flamingo|2022|4.6×10^11 個樣本OPT-175B|2022|1.8×10^11 個樣本UL2|2022|1.0×10^12 個樣本Gato|2022|5.2×10^11 個樣本SimCSE|2022|2.7×10^7 個樣本GPT-2 Medium (FlashAttention)|2022|1.0×10^10 個樣本Tranception|2022|4.8×10^10 個樣本CogVideo|2022|1.5×10^11 個樣本DITTO|2022|1.0×10^8 個樣本CoCa|2022|1.4×10^12 個樣本MetaLM|2022|6.5×10^11 個樣本Parti|2022|4.7×10^12 個樣本ProGen2-xlarge|2022|3.5×10^11 個樣本Minerva (540B)|2022|2.6×10^10 個樣本CodeT5-large|2022|1.1×10^10 個樣本NLLB|2022|3.0×10^11 個樣本BLOOM-176B|2022|3.8×10^11 個樣本ESM2-15B|2022|1.5×10^10 個樣本OmegaPLM|2022|1.3×10^12 個樣本AlexaTM 20B|2022|1.3×10^12 個樣本GLM-130B|2022|1.5×10^11 個樣本BlenderBot 3|2022|1.3×10^9 個樣本PaLI|2022|1.4×10^11 個樣本Whisper|2022|1.2×10^10 個樣本DiffDock|2022|4.4×10^6 個樣本GenSLM|2022|2.3×10^11 個樣本Flan-PaLM 540B|2022|1.4×10^9 個樣本LMSI-Palm|2022|1.9×10^6 個樣本U-PaLM (540B)|2022|1.3×10^9 個樣本Mogrifier RLSTM (WT2)|2022|2.7×10^6 個樣本eDiff-I|2022|1.6×10^12 個樣本mT0-13B|2022|2.0×10^10 個樣本InternImage|2022|8.4×10^10 個樣本EVA-01|2022|7.6×10^9 個樣本Galactica|2022|1.1×10^11 個樣本Fusion in Encoder|2022|9.6×10^5 個樣本ALM 1.0|2022|2.3×10^10 個樣本DiT-XL/2 + Discriminator Guidance|2022|3.3×10^8 個樣本Discriminator Guidance|2022|3.3×10^8 個樣本DeepNash|2022|2.1×10^12 個樣本Vega v2|2022|6.4×10^9 個樣本CaLM|2022|2.5×10^9 個樣本Hybrid H3-2.7B|2022|4.0×10^11 個樣本VALL-E|2023|7.7×10^10 個樣本DreamerV3|2023|1.6×10^9 個樣本Ankh_large|2023|1.4×10^10 個樣本Nucleotide Transformer|2023|3.0×10^11 個樣本DDPM-IP (CelebA)|2023|8.3×10^8 個樣本BLIP-2 (Q-Former)|2023|2.3×10^9 個樣本ProteinDT|2023|1.3×10^8 個樣本ViT-22B|2023|4.0×10^9 個樣本LLaMA-65B|2023|1.4×10^12 個樣本AudioGen|2023|2.3×10^11 個樣本Falcon-40B|2023|1.0×10^12 個樣本GPT-4 (Mar 2023)|2023|5.4×10^12 個樣本LEP-AD|2023|1.2×10^6 個樣本PanGu-Σ|2023|3.3×10^11 個樣本SigLIP 400M|2023|6.7×10^12 個樣本BloombergGPT|2023|5.7×10^11 個樣本VideoMAE V2|2023|1.2×10^9 個樣本Segment Anything Model|2023|1.1×10^9 個樣本Incoder-6.7B|2023|5.2×10^10 個樣本DINOv2|2023|3.6×10^10 個樣本Agile Soccer Robot|2023|3.1×10^9 個樣本PaLM 2|2023|3.6×10^12 個樣本StarCoder|2023|2.0×10^11 個樣本CoEdiT-xxl|2023|1.1×10^6 個樣本Med-PaLM 2|2023|1.6×10^7 個樣本CodeT5+|2023|5.2×10^10 個樣本ONE-PEACE|2023|4.9×10^11 個樣本Goat-7B|2023|4.4×10^7 個樣本MusicGen|2023|1.4×10^13 個樣本GPT-4 (Jun 2023)|2023|5.4×10^12 個樣本HyenaDNA|2023|3.0×10^9 個樣本InternLM|2023|1.6×10^12 個樣本Pangu-Weather|2023|2.5×10^13 個樣本xTrimoPGLM -100B|2023|2.8×10^11 個樣本Llama 2-70B|2023|2.0×10^12 個樣本Llama 2-7B|2023|2.0×10^12 個樣本AudioLM|2023|1.3×10^11 個樣本Qwen-VL|2023|5.0×10^11 個樣本PeptideBERT|2023|4.2×10^6 個樣本Jais|2023|4.0×10^11 個樣本Swift|2023|1.2×10^8 個樣本Falcon-180B|2023|3.5×10^12 個樣本AlphaMissense|2023|2.3×10^9 個樣本Amazon Titan|2023|4.0×10^12 個樣本Show-1|2023|1.6×10^14 個樣本FinGPT-13B|2023|7.7×10^4 個樣本RoseTTAFold All-Atom (RFAA)|2023|6.3×10^7 個樣本Ferret (13B)|2023|1.7×10^8 個樣本CODEFUSION (Python)|2023|4.4×10^6 個樣本ChatGLM3-6B|2023|1.4×10^12 個樣本Skywork-13B|2023|3.2×10^12 個樣本BLUUMI|2023|3.8×10^10 個樣本Grok-1|2023|6.2×10^12 個樣本Yi-34B|2023|3.1×10^12 個樣本RoFormer|2023|3.3×10^9 個樣本mPLUG-Owl2|2023|1.8×10^11 個樣本Nemotron-3-8B|2023|3.8×10^12 個樣本GNoME for crystal discovery|2023|6.9×10^4 個樣本Qwen-72B|2023|3.0×10^12 個樣本Mamba-24M (SC09)|2023|9.7×10^4 個樣本Llama Guard|2023|4.1×10^6 個樣本VILA-13B|2023|3.2×10^10 個樣本nekomata-14b|2023|6.6×10^10 個樣本Qwen1.5-72B|2024|3.0×10^12 個樣本Aya|2024|1.1×10^12 個樣本Aramco Metabrain AI|2024|7.0×10^12 個樣本DBRX|2024|1.2×10^13 個樣本ReALM|2024|1.3×10^11 個樣本Llama 3-70B|2024|1.5×10^13 個樣本VILA1.5-13B|2024|3.2×10^10 個樣本AlphaFold 3|2024|3.0×10^10 個樣本Yi-Large|2024|3.0×10^12 個樣本GLM-4 (0520)|2024|1.0×10^13 個樣本ALLaM adapted 70B|2024|6.0×10^11 個樣本Qwen2-72B|2024|7.0×10^12 個樣本Nemotron-4 340B|2024|9.0×10^12 個樣本DeepSeek-Coder-V2 236B|2024|3.2×10^12 個樣本ESM3 (98B)|2024|7.7×10^11 個樣本Llama 3.1-405B|2024|1.6×10^13 個樣本AFM-on-device|2024|7.6×10^12 個樣本AFM-server|2024|7.4×10^12 個樣本LLaVA-OV-72B|2024|3.8×10^10 個樣本Table Tennis Agent|2024|2.4×10^9 個樣本Qwen2.5-32B|2024|1.8×10^13 個樣本Qwen2.5-72B|2024|1.8×10^13 個樣本Telechat2-115B|2024|1.0×10^13 個樣本PixelDance|2024|1.1×10^14 個樣本Movie Gen Video|2024|3.4×10^9 個樣本NVLM-D 72B|2024|5.7×10^10 個樣本NVLM-H 72B|2024|1.3×10^11 個樣本NVLM-X 72B|2024|4.6×10^10 個樣本Doubao-pro|2024|8.4×10^12 個樣本Hunyuan-Large|2024|7.0×10^12 個樣本Llama 3.3 70B|2024|1.5×10^13 個樣本EXAONE 3.5 32B|2024|6.5×10^12 個樣本DeepSeek-V3|2024|1.5×10^13 個樣本DeepSeek-R1|2025|1.5×10^13 個樣本Doubao-1.5-pro|2025|9.0×10^12 個樣本Eurus-2-7B-PRIME|2025|8.3×10^5 個樣本Hunyuan-TurboS|2025|1.6×10^13 個樣本EXAONE Deep 32B|2025|1.2×10^10 個樣本DeepSeek-V3 (Mar 2025)|2025|1.5×10^13 個樣本Llama 4 Behemoth (preview)|2025|3.0×10^13 個樣本Llama 4 Maverick|2025|3.0×10^13 個樣本Llama 4 Scout|2025|3.0×10^13 個樣本Pangu Ultra|2025|1.3×10^13 個樣本Qwen3-235B-A22B|2025|3.6×10^13 個樣本Seed1.5-VL|2025|3.0×10^12 個樣本DeepSeek-R1 (May 2025)|2025|1.5×10^13 個樣本EXAONE Path 2.0|2025|1.4×10^5 個樣本Kimi K2|2025|1.6×10^13 個樣本EXAONE 4.0 (32B)|2025|1.4×10^13 個樣本Qwen3-Coder-480B-A35B|2025|7.5×10^12 個樣本Qwen3-235B-A22B (Jul 2025)|2025|3.6×10^13 個樣本Qwen3-235B-A22B-Thinking (Jul 2025)|2025|3.6×10^13 個樣本GLM-4.5|2025|2.3×10^13 個樣本LongCat-Flash|2025|2.3×10^13 個樣本Qwen3-Max|2025|3.6×10^13 個樣本AgentFounder-30B|2025|3.2×10^11 個樣本Qwen3-Omni-30B-A3B|2025|2.0×10^12 個樣本GLM-4.6|2025|2.3×10^13 個樣本Ling-1T|2025|2.0×10^13 個樣本Olmo 3|2025|5.5×10^12 個樣本K-EXAONE|2026|1.1×10^13 個樣本Solar Open 100B|2026|2.0×10^13 個樣本MiMo-V2.5-Pro|2026|2.7×10^13 個樣本Solar Open2 250B|2026|1.1×10^13 個樣本Inkling|2026|4.5×10^13 個樣本語言視覺多領域其他生物遊戲語音影像生成機器人
每個點為一個知名 AI 模型,縱軸=訓練資料集規模(對數軸,樣本 / token 數),按應用領域著色;共 662 個模型。

訓練算力紀錄:歷年重新整理前沿的模型

下表為在其釋出時重新整理「已知最高訓練算力」紀錄的模型(按算力升序重新整理),倒序展示最近 12 個紀錄。

年份模型領域訓練算力(petaFLOP)
2025Grok 4其他5.0×10¹¹
2025GPT-4.5其他3.8×10¹¹
2025Grok 3其他3.5×10¹¹
2023Gemini 1.0 Ultra其他5.0×10¹⁰
2023GPT-4 (Mar 2023)其他2.1×10¹⁰
2022Minerva (540B)其他2.7×10⁹
2022GPT-3.5 (davinci-002)其他2.6×10⁹
2021FLAN 137B其他2.1×10⁹
2021Jurassic-1-Jumbo其他3.7×10⁸
2020GPT-3 175B (davinci)其他3.1×10⁸
2020Meena其他1.1×10⁸
2019AlphaStar其他1.1×10⁸

延伸閱讀

資料來源:Epoch AI「Notable AI models」資料集(CC BY 4.0), 經 Our World in Data 整理。原始資料頁:訓練算力 ·引數量 ·訓練資料量。 資料於 2026-09-01 抓取,慢資料(約年度更新)、定期重新整理;各圖縱軸均為對數刻度。 本頁僅客觀呈現已公開資料,不預測、不構成任何投資建議。