Inference Providers
Active filters: quark
superbigtree/Mistral-Nemo-Instruct-2407-FP8
12B • Updated • 8
aigdat/BioMistral-7B_quantized_int4_float16
1B • Updated • 12
aigdat/omost-phi-3-mini-128k_quantized_int4_float16
0.6B • Updated • 10
superbigtree/Mistral-Nemo-Instruct-2407-FP8_aq
12B • Updated • 9
aigdat/Llama-3.2-1B-Instruct-awq-uint4-float16
0.4B • Updated • 11
aigdat/Llama-3.2-3B-Instruct-awq-uint4-float16
0.8B • Updated • 12
aigdat/Phi-3.5-mini-instruct-awq-uint4-float16
0.6B • Updated • 11
aigdat/DeepSeek-R1-Distill-Qwen-1.5B_quantized_int4_bfloat16
0.4B • Updated • 10
aigdat/Qwen3-0.6B_quantized_int4_float16
0.2B • Updated • 14
aigdat/Arch-Function-Chat-3B_quantized_int4_float16
0.7B • Updated • 12
aigdat/DeepCoder-14B-Preview_quantized_int4_float16
3B • Updated • 11
aigdat/Qwen2.5-Coder-1.5B-Instruct_quantized_int4_bfloat16
0.4B • Updated • 15
aigdat/Qwen2.5-Coder-7B-Instruct_quantized_int4_bfloat16
1B • Updated • 9
aigdat/Qwen2.5-3B-Instruct_quantized_int4_bfloat16
0.7B • Updated • 10
aigdat/Qwen2.5-Coder-32B-Instruct_quantized_int4_bfloat16
5B • Updated • 11
aigdat/Llama-xLAM-2-8b-fc-r_quantized_int4_bfloat16
2B • Updated • 142
fxmarty/qwen_1.5-moe-a2.7b-mxfp4
8B • Updated • 7.05k
amd/Llama-3.3-70B-Instruct-MXFP4-Preview
38B • Updated • 2.11k
• 2
fxmarty/deepseek_r1_3_layers_mxfp4
8B • Updated • 90
• 1
fxmarty/Llama-4-Scout-17B-16E-Instruct-2-layers-mxfp4
5B • Updated • 15
• 1
371B • Updated • 68.9k
• 5
mohitsha/Llama-2-7b-hf-w_mx_fp4_per_group_sym
4B • Updated • 9
amd/Llama-3.1-405B-Instruct-MXFP4-Preview
218B • Updated • 862
• 1
amd/DeepSeek-R1-MXFP4-ASQ
363B • Updated • 16.6k
• 1
haoyang-amd/qwen1.5-0.5B-ptpc
0.5B • Updated • 8
amd/DeepSeek-R1-0528-MXFP4
356B • Updated • 9.14k
• 2
fxmarty/Llama-3.1-70B-Instruct-2-layers-mxfp6
3B • Updated • 10.9k
fxmarty/qwen1.5_moe_a2.7b_chat_w_fp4_a_fp6_e2m3
8B • Updated • 6.79k
fxmarty/qwen1.5_moe_a2.7b_chat_w_fp6_e2m3_a_fp6_e2m3
11B • Updated • 15
fxmarty/qwen1.5_moe_a2.7b_chat_w_fp6_e3m2_a_fp6_e3m2
11B • Updated • 6.73k