-
Attention Is All You Need
Paper • 1706.03762 • Published • 134 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 112M • • 2.72k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 13.6M • 3.39k
Collections
Discover the best community collections!
Collections including paper arxiv:2302.04761
-
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 2.01M • • 765 -
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.92M • • 6.45k -
Toolformer: Language Models Can Teach Themselves to Use Tools
Paper • 2302.04761 • Published • 12 -
ReAct: Synergizing Reasoning and Acting in Language Models
Paper • 2210.03629 • Published • 36
-
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Paper • 2307.16789 • Published • 102 -
Tool Documentation Enables Zero-Shot Tool-Usage with Large Language Models
Paper • 2308.00675 • Published • 37 -
Toolformer: Language Models Can Teach Themselves to Use Tools
Paper • 2302.04761 • Published • 12 -
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Paper • 2305.18752 • Published • 5
-
Attention Is All You Need
Paper • 1706.03762 • Published • 134 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 10 -
Training Compute-Optimal Large Language Models
Paper • 2203.15556 • Published • 12 -
Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT
Paper • 2210.04186 • Published
-
Neural Machine Translation by Jointly Learning to Align and Translate
Paper • 1409.0473 • Published • 7 -
Attention Is All You Need
Paper • 1706.03762 • Published • 134 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 32 -
Hierarchical Reasoning Model
Paper • 2506.21734 • Published • 54
-
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Paper • 2503.12605 • Published • 35 -
AppAgentX: Evolving GUI Agents as Proficient Smartphone Users
Paper • 2503.02268 • Published • 11 -
Machine Learning Operations (MLOps): Overview, Definition, and Architecture
Paper • 2205.02302 • Published • 1 -
Beyond Browsing: API-Based Web Agents
Paper • 2410.16464 • Published • 2
-
Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning
Paper • 2211.04325 • Published • 1 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 32 -
On the Opportunities and Risks of Foundation Models
Paper • 2108.07258 • Published • 2 -
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Paper • 2204.07705 • Published • 2
-
Attention Is All You Need
Paper • 1706.03762 • Published • 134 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 112M • • 2.72k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 13.6M • 3.39k
-
Attention Is All You Need
Paper • 1706.03762 • Published • 134 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 10 -
Training Compute-Optimal Large Language Models
Paper • 2203.15556 • Published • 12 -
Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT
Paper • 2210.04186 • Published
-
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 2.01M • • 765 -
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.92M • • 6.45k -
Toolformer: Language Models Can Teach Themselves to Use Tools
Paper • 2302.04761 • Published • 12 -
ReAct: Synergizing Reasoning and Acting in Language Models
Paper • 2210.03629 • Published • 36
-
Neural Machine Translation by Jointly Learning to Align and Translate
Paper • 1409.0473 • Published • 7 -
Attention Is All You Need
Paper • 1706.03762 • Published • 134 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 32 -
Hierarchical Reasoning Model
Paper • 2506.21734 • Published • 54
-
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Paper • 2503.12605 • Published • 35 -
AppAgentX: Evolving GUI Agents as Proficient Smartphone Users
Paper • 2503.02268 • Published • 11 -
Machine Learning Operations (MLOps): Overview, Definition, and Architecture
Paper • 2205.02302 • Published • 1 -
Beyond Browsing: API-Based Web Agents
Paper • 2410.16464 • Published • 2
-
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Paper • 2307.16789 • Published • 102 -
Tool Documentation Enables Zero-Shot Tool-Usage with Large Language Models
Paper • 2308.00675 • Published • 37 -
Toolformer: Language Models Can Teach Themselves to Use Tools
Paper • 2302.04761 • Published • 12 -
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Paper • 2305.18752 • Published • 5
-
Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning
Paper • 2211.04325 • Published • 1 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 32 -
On the Opportunities and Risks of Foundation Models
Paper • 2108.07258 • Published • 2 -
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Paper • 2204.07705 • Published • 2