meta-llama/meta-llama/Llama-3.3-70B-Instruct-Turbo
40B-100B持续维护
任务类型对话模型
参数规模70B
上下文长度131072
开源协议暂无可靠数据
商用情况暂无可靠数据
推荐显存暂无可靠数据
主要语言暂无可靠数据
本地部署未标注
模型介绍
本条目的官方说明为英文,以下内容保留原文,关键章节标题已中文标注。
meta-llama/Llama-3.3-70B-Instruct-Turbo
Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.
- 上下文长度:131,072 tokens
- 最大输出:131,072 tokens
- 标签:chat
- 定价(每百万 token,USD):输入 $0.1000 / 输出 $0.3200
- 提供方:meta-llama
- 来源:DeepInfra
适合场景
暂无可靠数据。
已知限制
- 本站未登记确切发布日期,版本时间线可能不完整。
- 本页评测与硬件数据为区间估算,未做统一环境实测,不能替代官方评测报告。
评测
暂无可靠数据。本站只收录标注了「评测来源、模型版本、是否官方数据、测试日期」的成绩, 不混合不同设置下的分数,也不把单一 Benchmark 解释为整体能力。