{"data":[{"id":"anthropic/claude-opus-5","name":"Anthropic: Claude Opus 5","created":1786025346,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1000000,"max_output_length":1000000,"pricing":{"prompt":"0.000005","completion":"0.000025"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["anthropic"],"description":"Claude Opus 5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis.","metadata":{}},{"id":"openai/gpt-5-nano","name":"OpenAI: GPT-5 Nano","created":1754616202,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":400000,"max_output_length":400000,"pricing":{"prompt":"0.00000005","completion":"0.0000004"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger counterparts, it retains key instruction-following and safety features. It is the successor to GPT-4.1-nano and offers a lightweight option for cost-sensitive or real-time applications.","metadata":{}},{"id":"openai/o3","name":"OpenAI: o3","created":1744852257,"input_modalities":["image","text","file"],"output_modalities":["text"],"context_length":200000,"max_output_length":100000,"pricing":{"prompt":"0.000002","completion":"0.000008"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following. Use it to think through multi-step problems that involve analysis across text, code, and images.","metadata":{}},{"id":"deepseek/deepseek-v4-flash-0731","name":"DeepSeek: DeepSeek V4 Flash 0731","created":1785805281,"input_modalities":["text"],"output_modalities":["text"],"context_length":1048576,"max_output_length":393216,"pricing":{"prompt":"0.00000044","completion":"0.00000132","input_cache_read":"0.000000028"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"deepseek-ai/DeepSeek-V4-Flash-0731","is_tee":true,"providers":["chutes"],"description":"DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.","metadata":{}},{"id":"openai/gpt-5","name":"OpenAI: GPT-5","created":1754616213,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":400000,"max_output_length":128000,"pricing":{"prompt":"0.00000125","completion":"0.00001"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like \"think hard about this.\" Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.","metadata":{}},{"id":"openai/gpt-5-mini","name":"OpenAI: GPT-5 Mini","created":1754616207,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":400000,"max_output_length":128000,"pricing":{"prompt":"0.00000025","completion":"0.000002"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost. GPT-5 Mini is the successor to OpenAI's o4-mini model.","metadata":{}},{"id":"google/gemini-3.1-pro-preview","name":"Google: Gemini 3.1 Pro Preview","created":1786025372,"input_modalities":["audio","file","image","text","video"],"output_modalities":["text"],"context_length":1048576,"max_output_length":1048576,"pricing":{"prompt":"0.000004","completion":"0.000018"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["google"],"description":"Gemini 3.1 Pro Preview is Google's frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation.","metadata":{}},{"id":"openai/gpt-4.1-mini","name":"OpenAI: GPT-4.1 Mini","created":1744680181,"input_modalities":["image","text","file"],"output_modalities":["text"],"context_length":1047576,"max_output_length":32768,"pricing":{"prompt":"0.0000004","completion":"0.0000016"},"supported_parameters":["max_completion_tokens","max_tokens","response_format","seed","structured_outputs","temperature","tool_choice","tools","top_p"],"supported_sampling_parameters":["max_tokens","seed","temperature","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard instruction evals, 35.8% on MultiChallenge, and 84.1% on IFEval. Mini also shows strong coding ability (e.g., 31.6% on Aider’s polyglot diff benchmark) and vision understanding, making it suitable for interactive applications with tight performance constraints.","metadata":{}},{"id":"openai/gpt-4.1","name":"OpenAI: GPT-4.1","created":1744680185,"input_modalities":["image","text","file"],"output_modalities":["text"],"context_length":1047576,"max_output_length":1047576,"pricing":{"prompt":"0.000002","completion":"0.000008"},"supported_parameters":["max_completion_tokens","max_tokens","response_format","seed","structured_outputs","temperature","tool_choice","tools","top_p"],"supported_sampling_parameters":["max_tokens","seed","temperature","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and GPT-4.5 across coding (54.6% SWE-bench Verified), instruction compliance (87.4% IFEval), and multimodal understanding benchmarks. It is tuned for precise code diffs, agent reliability, and high recall in large document contexts, making it ideal for agents, IDE tooling, and enterprise knowledge retrieval.","metadata":{}},{"id":"openai/gpt-4.1-nano","name":"OpenAI: GPT-4.1 Nano","created":1744680169,"input_modalities":["image","text","file"],"output_modalities":["text"],"context_length":1047576,"max_output_length":32768,"pricing":{"prompt":"0.0000001","completion":"0.0000004"},"supported_parameters":["max_completion_tokens","max_tokens","response_format","seed","structured_outputs","temperature","tool_choice","tools","top_p"],"supported_sampling_parameters":["max_tokens","seed","temperature","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion.","metadata":{}},{"id":"openai/gpt-4o","name":"OpenAI: GPT-4o","created":1715587200,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":128000,"max_output_length":16384,"pricing":{"prompt":"0.0000025","completion":"0.00001"},"supported_parameters":["frequency_penalty","logit_bias","logprobs","max_completion_tokens","max_tokens","presence_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_logprobs","top_p","web_search_options"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","presence_penalty","seed","stop","temperature","top_p"],"supported_features":["json_mode","logprobs","structured_outputs","tools","web_search"],"is_tee":false,"providers":["openai"],"description":"GPT-4o (\"o\" for \"omni\") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twice as fast and 50% more cost-effective. GPT-4o also offers improved performance in processing non-English languages and enhanced visual capabilities. For benchmarking against other models, it was briefly called \"im-also-a-good-gpt2-chatbot\" #multimodal","metadata":{}},{"id":"openai/gpt-3.5-turbo","name":"OpenAI: GPT-3.5 Turbo","created":1685260800,"input_modalities":["text"],"output_modalities":["text"],"context_length":16385,"max_output_length":4096,"pricing":{"prompt":"0.0000005","completion":"0.0000015"},"supported_parameters":["frequency_penalty","logit_bias","logprobs","max_tokens","presence_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","presence_penalty","seed","stop","temperature","top_p"],"supported_features":["json_mode","logprobs","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.","metadata":{}},{"id":"openai/gpt-4","name":"OpenAI: GPT-4","created":1685260800,"input_modalities":["text"],"output_modalities":["text"],"context_length":8191,"max_output_length":4096,"pricing":{"prompt":"0.00003","completion":"0.00006"},"supported_parameters":["frequency_penalty","logit_bias","logprobs","max_completion_tokens","max_tokens","presence_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","presence_penalty","seed","stop","temperature","top_p"],"supported_features":["json_mode","logprobs","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning capabilities. Training data: up to Sep 2021.","metadata":{}},{"id":"openai/gpt-5.6-luna","name":"OpenAI: GPT-5.6 Luna","created":1786025428,"input_modalities":["file","image","text"],"output_modalities":["text"],"context_length":1050000,"max_output_length":1050000,"pricing":{"prompt":"0.0000004","completion":"0.0000018"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["openai"],"description":"GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning.","metadata":{}},{"id":"google/gemini-3.6-flash","name":"Google: Gemini 3.6 Flash","created":1786025381,"input_modalities":["text","image","video","file","audio"],"output_modalities":["text"],"context_length":1048576,"max_output_length":1048576,"pricing":{"prompt":"0.0000015","completion":"0.0000075"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["google"],"description":"Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits.","metadata":{}},{"id":"google/gemini-3.5-flash-lite","name":"Google: Gemini 3.5 Flash Lite","created":1786025389,"input_modalities":["text","image","video","file","audio"],"output_modalities":["text"],"context_length":1048576,"max_output_length":1048576,"pricing":{"prompt":"0.0000003","completion":"0.0000025"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["google"],"description":"Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.","metadata":{}},{"id":"openai/gpt-5.6-sol","name":"OpenAI: GPT-5.6 Sol","created":1786025410,"input_modalities":["file","image","text"],"output_modalities":["text"],"context_length":1050000,"max_output_length":1050000,"pricing":{"prompt":"0.00001","completion":"0.000045"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["openai"],"description":"GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks.","metadata":{}},{"id":"x-ai/grok-code-fast-1","name":"xAI: Grok Code Fast 1","created":1765626614,"input_modalities":["text"],"output_modalities":["text"],"context_length":256000,"max_output_length":256000,"pricing":{"prompt":"0.0000002","completion":"0.0000015"},"supported_parameters":["include_reasoning","logprobs","max_tokens","reasoning","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_logprobs","top_p"],"supported_sampling_parameters":["max_tokens","seed","stop","temperature","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["x-ai"],"description":"Grok Code Fast 1 is a speedy and economical reasoning model that excels at agentic coding. With reasoning traces visible in the response, developers can steer Grok Code for high-quality work flows.","metadata":{}},{"id":"openai/gpt-5.1","name":"OpenAI: GPT-5.1","created":1763127098,"input_modalities":["image","text","file"],"output_modalities":["text"],"context_length":400000,"max_output_length":128000,"pricing":{"prompt":"0.00000125","completion":"0.00001"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning to allocate computation dynamically, responding quickly to simple queries while spending more depth on complex tasks. The model produces clearer, more grounded explanations with reduced jargon, making it easier to follow even on technical or multi-step problems. Built for broad task coverage, GPT-5.1 delivers consistent gains across math, coding, and structured analysis workloads, with more coherent long-form answers and improved tool-use reliability. It also features refined conversational alignment, enabling warmer, more intuitive responses without compromising precision. GPT-5.1 serves as the primary full-capability successor to GPT-5","metadata":{}},{"id":"openai/gpt-5.6-terra","name":"OpenAI: GPT-5.6 Terra","created":1786025419,"input_modalities":["file","image","text"],"output_modalities":["text"],"context_length":1050000,"max_output_length":1050000,"pricing":{"prompt":"0.000004","completion":"0.000018"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["openai"],"description":"GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic work.","metadata":{}},{"id":"meta-llama/llama-3.3-70b-instruct","name":"Meta: Llama 3.3 70B Instruct","created":1764320401,"input_modalities":["text"],"output_modalities":["text"],"context_length":131072,"max_output_length":16384,"pricing":{"prompt":"0.000002","completion":"0.000002"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"hugging_face_id":"meta-llama/Llama-3.3-70B-Instruct","is_tee":true,"providers":["tinfoil"],"description":"The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.","metadata":{}},{"id":"moonshotai/kimi-k2.7-code","name":"MoonshotAI: Kimi K2.7 Code","created":1786723082,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000095","completion":"0.000004","input_cache_read":"0.00000019"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"moonshotai/Kimi-K2.7-Code","is_tee":false,"providers":["redpill"],"description":"MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...","metadata":{}},{"id":"tencent/hy3","name":"Tencent: Hy3","created":1787324111,"input_modalities":["text"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000015","completion":"0.00000064","input_cache_read":"0.00000004"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","max_completion_tokens","max_tokens","min_p","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"hugging_face_id":"tencent/Hy3","is_tee":false,"providers":["redpill"],"description":"Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort.","metadata":{}},{"id":"deepseek/deepseek-v4-pro-0813","name":"DeepSeek: DeepSeek V4 Pro 0813","created":1786871243,"input_modalities":["text"],"output_modalities":["text"],"context_length":1048576,"max_output_length":393216,"pricing":{"prompt":"0.00000145","completion":"0.00000436","input_cache_read":"0.00000015"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"deepseek-ai/DeepSeek-V4-Pro-0813","is_tee":false,"providers":["redpill"],"description":"DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.","metadata":{}},{"id":"google/gemini-2.5-pro","name":"Google: Gemini 2.5 Pro","created":1750198344,"input_modalities":["text","image","file","audio","video"],"output_modalities":["text"],"context_length":1048576,"max_output_length":65536,"pricing":{"prompt":"0.0000025","completion":"0.000015"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_p"],"supported_sampling_parameters":["max_tokens","seed","stop","temperature","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["google"],"description":"Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy and nuanced context handling. Gemini 2.5 Pro achieves top-tier performance on multiple benchmarks, including first-place positioning on the LMArena leaderboard, reflecting superior human-preference alignment and complex problem-solving abilities.","metadata":{}},{"id":"qwen/qwen3.8-27b","name":"Qwen: Qwen3.8 27B","created":1787560779,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000024","completion":"0.0000025","input_cache_read":"0.00000005"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3.8-27B","is_tee":true,"providers":["chutes","phala"],"description":"Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. Served on Phala in a TDX-attested enclave.","metadata":{}},{"id":"anthropic/claude-haiku-4.5","name":"Anthropic: Claude Haiku 4.5","created":1760576438,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":200000,"max_output_length":64000,"pricing":{"prompt":"0.000001","completion":"0.000005"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["max_tokens","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["anthropic"],"description":"Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance across reasoning, coding, and computer-use tasks, Haiku 4.5 brings frontier-level capability to real-time and high-volume applications. It introduces extended thinking to the Haiku line; enabling controllable reasoning depth, summarized or interleaved thought output, and tool-assisted workflows with full support for coding, bash, web search, and computer-use tools. Scoring >73% on SWE-bench Verified, Haiku 4.5 ranks among the world’s best coding models while maintaining exceptional responsiveness for sub-agents, parallelized execution, and scaled deployment.","metadata":{}},{"id":"meta/muse-glimmer-30b","name":"Meta: Muse Glimmer 30B","created":1786593197,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":131072,"max_output_length":131072,"pricing":{"prompt":"0.0000003","completion":"0.0000011","input_cache_read":"0.00000004"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"meta-models/Muse-Glimmer-30B","is_tee":true,"providers":["phala"],"description":"Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon agentic and coding workflows, with multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across more than 100 languages.","metadata":{}},{"id":"anthropic/claude-sonnet-4.6","name":"Anthropic: Claude Sonnet 4.6","created":1772077178,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1000000,"max_output_length":128000,"pricing":{"prompt":"0.000003","completion":"0.000015"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p","verbosity"],"supported_sampling_parameters":["max_tokens","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["anthropic"],"description":"Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation.","metadata":{}},{"id":"anthropic/claude-opus-4.6","name":"Anthropic: Claude Opus 4.6","created":1770361489,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1000000,"max_output_length":128000,"pricing":{"prompt":"0.00001","completion":"0.0000375"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p","verbosity"],"supported_sampling_parameters":["max_tokens","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["anthropic"],"description":"Opus 4.6 is Anthropic's strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time.","metadata":{}},{"id":"z-ai/glm-5.2","name":"Z.ai: GLM 5.2","created":1781637446,"input_modalities":["text"],"output_modalities":["text"],"context_length":1048576,"max_output_length":131072,"pricing":{"prompt":"0.00000126","completion":"0.000003","input_cache_read":"0.00000022"},"supported_parameters":["frequency_penalty","include_reasoning","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"quantization":"fp8","hugging_face_id":"zai-org/GLM-5.2","is_tee":true,"providers":["chutes"],"description":"GLM-5.2 is Z.ai's flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context and execute long-running tasks more reliably. Served as a text-only TEE deployment via Phala.","metadata":{}},{"id":"openai/gpt-5.4","name":"OpenAI: GPT-5.4","created":1778224692,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1050000,"max_output_length":128000,"pricing":{"prompt":"0.0000025","completion":"0.000015"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5.4 is OpenAI's latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...","metadata":{}},{"id":"openai/gpt-5.5","name":"OpenAI: GPT-5.5","created":1778224693,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1050000,"max_output_length":128000,"pricing":{"prompt":"0.000005","completion":"0.00003"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5.5 is OpenAI's frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...","metadata":{}},{"id":"qwen/qwen3.5-397b-a17b","name":"Qwen: Qwen3.5 397B A17B","created":1772249193,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000055","completion":"0.0000035","input_cache_read":"0.000000225"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3.5-397B-A17B","is_tee":true,"providers":["chutes"],"description":"The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent.","metadata":{}},{"id":"qwen/qwen3.5-27b","name":"Qwen: Qwen3.5-27B","created":1773395704,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":65536,"pricing":{"prompt":"0.0000003","completion":"0.0000024","input_cache_read":"0.00000015"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3.5-27B","is_tee":true,"providers":["chutes"],"description":"The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.","metadata":{}},{"id":"openai/gpt-5.4-mini","name":"OpenAI: GPT-5.4 Mini","created":1778224698,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":400000,"max_output_length":128000,"pricing":{"prompt":"0.00000075","completion":"0.0000045"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...","metadata":{}},{"id":"openai/gpt-4o-mini","name":"OpenAI: GPT-4o-mini","created":1721289600,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":128000,"max_output_length":16384,"pricing":{"prompt":"0.00000015","completion":"0.0000006"},"supported_parameters":["frequency_penalty","logit_bias","logprobs","max_completion_tokens","max_tokens","presence_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_logprobs","top_p","web_search_options"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","presence_penalty","seed","stop","temperature","top_p"],"supported_features":["json_mode","logprobs","structured_outputs","tools","web_search"],"is_tee":false,"providers":["openai"],"description":"GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than GPT-3.5 Turbo. It maintains SOTA intelligence, while being significantly more cost-effective. GPT-4o mini achieves an 82% score on MMLU and presently ranks higher than GPT-4 on chat preferences common leaderboards. Check out the launch announcement to learn more. #multimodal","metadata":{}},{"id":"z-ai/glm-5.3-flash","name":"Z.ai: GLM 5.3 Flash","created":1787873595,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":1048576,"max_output_length":131072,"pricing":{"prompt":"0.00000015","completion":"0.00000050","input_cache_read":"0.00000003"},"supported_parameters":["frequency_penalty","include_reasoning","logprobs","max_completion_tokens","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"quantization":"fp8","hugging_face_id":"zai-org/GLM-5.3-Flash","is_tee":true,"providers":["near-ai"],"description":"GLM-5.3-Flash is Z.ai's natively multimodal 320B MoE model with 18B active parameters, designed for efficient coding, long-horizon agent tasks, visual understanding, and long-context inference. Served as a TEE deployment via Phala.","metadata":{}},{"id":"openai/gpt-5.2","name":"OpenAI: GPT-5.2","created":1765626796,"input_modalities":["file","image","text"],"output_modalities":["text"],"context_length":400000,"max_output_length":128000,"pricing":{"prompt":"0.00000175","completion":"0.000014"},"supported_parameters":["include_reasoning","max_completion_tokens","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly to simple queries while spending more depth on complex tasks. Built for broad task coverage, GPT-5.2 delivers consistent gains across math, coding, sciende, and tool calling workloads, with more coherent long-form answers and improved tool-use reliability.","metadata":{}},{"id":"nvidia/nemotron-3.5-lightning","name":"NVIDIA: Nemotron 3.5 Lightning","created":1789120395,"input_modalities":["text"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000007","completion":"0.00000020","input_cache_read":"0.00000004"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16","is_tee":true,"providers":["phala"],"description":"NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...","metadata":{}},{"id":"z-ai/glm-5.3","name":"Z.ai: GLM 5.3","created":1788145378,"input_modalities":["text"],"output_modalities":["text"],"context_length":1048576,"max_output_length":131072,"pricing":{"prompt":"0.0000014","completion":"0.0000044","input_cache_read":"0.00000026"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_completion_tokens","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"zai-org/GLM-5.3","is_tee":true,"providers":["phala"],"description":"GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window. Served as a text-only TEE deployment via Phala.","metadata":{}},{"id":"qwen/qwen3.6-27b","name":"Qwen: Qwen3.6 27B","created":1780549917,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":262140,"pricing":{"prompt":"0.00000032","completion":"0.0000027","input_cache_read":"0.00000015"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"Qwen/Qwen3.6-27B","is_tee":true,"providers":["chutes","phala"],"description":"Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities accepting text and image inputs, a configurable thinking/reasoning mode, and a native 262K context window. Served as a TEE deployment via Chutes.","metadata":{}},{"id":"x-ai/grok-4.1-fast","name":"xAI: Grok 4.1 Fast","created":1765626615,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":2000000,"max_output_length":2000000,"pricing":{"prompt":"0.0000002","completion":"0.0000005"},"supported_parameters":["include_reasoning","logprobs","max_tokens","reasoning","response_format","seed","structured_outputs","temperature","tool_choice","tools","top_logprobs","top_p"],"supported_sampling_parameters":["max_tokens","seed","temperature","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["x-ai"],"description":"Grok 4.1 Fast is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window.","metadata":{}},{"id":"x-ai/grok-4","name":"xAI: Grok 4","created":1765626614,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":256000,"max_output_length":256000,"pricing":{"prompt":"0.000003","completion":"0.000015"},"supported_parameters":["include_reasoning","logprobs","max_tokens","reasoning","response_format","seed","structured_outputs","temperature","tool_choice","tools","top_logprobs","top_p"],"supported_sampling_parameters":["max_tokens","seed","temperature","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["x-ai"],"description":"Grok 4 is xAI's latest reasoning model with a 256k context window. It supports parallel tool calling, structured outputs, and both image and text inputs. Note that reasoning is not exposed, reasoning cannot be disabled, and the reasoning effort cannot be specified. Pricing increases once the total tokens in a given request is greater than 128k tokens. See more details on the [xAI docs](https://docs.x.ai/docs/models/grok-4-0709)","metadata":{}},{"id":"anthropic/claude-sonnet-4.5","name":"Anthropic: Claude Sonnet 4.5","created":1759190476,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1000000,"max_output_length":64000,"pricing":{"prompt":"0.000003","completion":"0.000015"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["max_tokens","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["anthropic"],"description":"Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements across system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking. Sonnet 4.5 also introduces stronger agentic capabilities, including improved tool orchestration, speculative parallel execution, and more efficient context and memory management. With enhanced context tracking and awareness of token usage across tool calls, it is particularly well-suited for multi-context and long-running workflows. Use cases span software engineering, cybersecurity, financial analysis, research agents, and other domains requiring sustained reasoning and tool use.","metadata":{}},{"id":"google/gemini-2.5-flash-lite","name":"Google: Gemini 2.5 Flash Lite","created":1753229076,"input_modalities":["text","image","file","audio","video"],"output_modalities":["text"],"context_length":1048576,"max_output_length":65535,"pricing":{"prompt":"0.0000001","completion":"0.0000004"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_p"],"supported_sampling_parameters":["max_tokens","seed","stop","temperature","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["google"],"description":"Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, \"thinking\" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter to selectively trade off cost for intelligence.","metadata":{}},{"id":"phala/qwen3.8-27b-uncensored","name":"Phala: Qwen3.8 27B Uncensored (Aggressive)","created":1789549583,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.0000003","completion":"0.0000015"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"is_tee":true,"providers":["phala"],"description":"HauhauCS Qwen3.8-27B Uncensored Aggressive, using the pinned Q8_K_P GGUF and BF16 vision projector. Served with SGLang on a Phala TDX-attested H200, with text and image input and a 262144-token context.","metadata":{}},{"id":"deepseek/deepseek-v4.1-flash","name":"DeepSeek: DeepSeek V4.1 Flash","created":1789403150,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":1048576,"max_output_length":393216,"pricing":{"prompt":"0.000000345","completion":"0.00000138","input_cache_read":"0.0000000069"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"deepseek-ai/DeepSeek-V4.1-Flash","is_tee":false,"providers":["redpill"],"description":"DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...","metadata":{}},{"id":"google/gemini-2.5-flash","name":"Google: Gemini 2.5 Flash","created":1750201288,"input_modalities":["file","image","text","audio","video"],"output_modalities":["text"],"context_length":1048576,"max_output_length":65535,"pricing":{"prompt":"0.0000003","completion":"0.0000025"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_p"],"supported_sampling_parameters":["max_tokens","seed","stop","temperature","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["google"],"description":"Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in \"thinking\" capabilities, enabling it to provide responses with greater accuracy and nuanced context handling. Additionally, Gemini 2.5 Flash is configurable through the \"max tokens for reasoning\" parameter, as described in the documentation (","metadata":{}},{"id":"nousresearch/hermes-3-llama-3.1-405b","name":"Nous: Hermes 3 405B Instruct","created":1723795200,"input_modalities":["text"],"output_modalities":["text"],"context_length":131072,"max_output_length":16384,"pricing":{"prompt":"0.000001","completion":"0.000001"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","structured_outputs"],"hugging_face_id":"NousResearch/Hermes-3-Llama-3.1-405B","is_tee":false,"providers":["openrouter"],"description":"Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the board. Hermes 3 405B is a frontier-level, full-parameter finetune of the Llama-3.1 405B foundation model, focused on aligning LLMs to the user, with powerful steering capabilities and control given to the end user. The Hermes 3 series builds and expands on the Hermes 2 set of capabilities, including more powerful and reliable function calling and structured output capabilities, generalist assistant capabilities, and improved code generation skills. Hermes 3 is competitive, if not superior, to Llama-3.1 Instruct models at general capabilities, with varying strengths and weaknesses attributable between the two.","metadata":{}},{"id":"openai/gpt-oss-120b","name":"OpenAI: GPT OSS 120B","created":1759532710,"input_modalities":["text"],"output_modalities":["text"],"context_length":131072,"max_output_length":131072,"pricing":{"prompt":"0.00000015","completion":"0.0000006"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"openai/gpt-oss-120b","is_tee":true,"providers":["secretai","tinfoil"],"description":"gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.","metadata":{}},{"id":"openai/o4-mini","name":"OpenAI: o4 Mini","created":1744849742,"input_modalities":["image","text","file"],"output_modalities":["text"],"context_length":200000,"max_output_length":100000,"pricing":{"prompt":"0.0000011","completion":"0.0000044"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","seed","structured_outputs","tool_choice","tools"],"supported_sampling_parameters":["max_tokens","seed"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["openai"],"description":"OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning and coding performance across benchmarks like AIME (99.5% with Python) and SWE-bench, outperforming its predecessor o3-mini and even approaching o3 in some domains. Despite its smaller size, o4-mini exhibits high accuracy in STEM tasks, visual problem solving (e.g., MathVista, MMMU), and code editing. It is especially well-suited for high-throughput scenarios where latency or cost is critical. Thanks to its efficient architecture and refined reinforcement learning training, o4-mini can chain tools, generate structured outputs, and solve multi-step tasks with minimal delay—often in under a minute.","metadata":{}},{"id":"phala/gemma-4-26b-a4b-uncensored","name":"Phala: Gemma-4 26B-A4B Uncensored (Heretic)","created":1779562961,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000015","completion":"0.0000007"},"supported_parameters":["frequency_penalty","logit_bias","logprobs","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","structured_outputs","tools"],"is_tee":true,"providers":["phala"],"description":"Uncensored \"Heretic\" variant of google/gemma-4-26B-A4B-it created using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method and row-norm preservation. Refusals drop from 100/100 to 11/100 with KL divergence 0.0499 vs the base model. The base Gemma 4 26B A4B is a Mixture-of-Experts model with 25.2B total / 3.8B active parameters (8 active / 128 total experts), 30-layer transformer with hybrid local sliding (1024) + global attention, supporting a 256K context window. Natively multimodal (text + images, variable aspect ratios). Strong on coding, reasoning, function calling, with native system prompt support across 35+ languages. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; vLLM-compatible FP8-Static quantization by cloud19 (router excluded from quantization).","metadata":{}},{"id":"anthropic/claude-opus-4.7","name":"Anthropic: Claude Opus 4.7","created":1778224978,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1000000,"max_output_length":128000,"pricing":{"prompt":"0.000005","completion":"0.000025"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","stop","structured_outputs","tool_choice","tools","verbosity"],"supported_sampling_parameters":["max_tokens","stop"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["anthropic"],"description":"Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...","metadata":{}},{"id":"moonshotai/kimi-k2.6","name":"MoonshotAI: Kimi K2.6","created":1776747825,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000109","completion":"0.0000046","input_cache_read":"0.00000037"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"moonshotai/Kimi-K2.6","is_tee":true,"providers":["chutes"],"description":"Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and demonstrates strong performance in agentic workflows.","metadata":{}},{"id":"google/gemma-4-31b-it","name":"Google: Gemma 4 31B","created":1779762858,"input_modalities":["image","text","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000015","completion":"0.00000046","input_cache_read":"0.000000075"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_a","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_a","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"hugging_face_id":"google/gemma-4-31B-it","is_tee":true,"providers":["chutes","tinfoil"],"description":"Gemma 4 31B Instruct is Google DeepMind's 30.7B dense model. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and strong multilingual performance. Served as a text-only TEE deployment via NEAR AI.","metadata":{}},{"id":"qwen/qwen3-32b","name":"Qwen: Qwen3 32B","created":1779762873,"input_modalities":["text"],"output_modalities":["text"],"context_length":40960,"max_output_length":16384,"pricing":{"prompt":"0.00000012","completion":"0.0000005","input_cache_read":"0.000000052"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"Qwen/Qwen3-32B","is_tee":true,"providers":["chutes"],"description":"Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a thinking mode for complex tasks and a standard mode for general dialogue. Served as a TEE deployment via Chutes.","metadata":{}},{"id":"qwen/qwen3.6-35b-a3b","name":"Qwen: Qwen3.6 35B A3B","created":1779762842,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.0000002","completion":"0.00000127"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"Qwen/Qwen3.6-35B-A3B","is_tee":true,"providers":["near-ai"],"description":"Qwen3.6-35B-A3B is an open-weight model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated attention. Served as a text-only TEE deployment via NEAR AI.","metadata":{}},{"id":"anthropic/claude-opus-4.8","name":"Anthropic: Claude Opus 4.8","created":1780068415,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1000000,"max_output_length":128000,"pricing":{"prompt":"0.000005","completion":"0.000025"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","stop","structured_outputs","tool_choice","tools","verbosity"],"supported_sampling_parameters":["max_tokens","stop"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["anthropic"],"description":"Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, includes reasoning support, and has a 1M-token context window.","metadata":{}},{"id":"anthropic/claude-sonnet-5","name":"Anthropic: Claude Sonnet 5","created":1783007094,"input_modalities":["text","image","file"],"output_modalities":["text"],"context_length":1000000,"max_output_length":128000,"pricing":{"prompt":"0.000003","completion":"0.000015"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"is_tee":false,"providers":["anthropic"],"description":"Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max).","metadata":{}},{"id":"z-ai/glm-5.1","name":"Z.ai: GLM 5.1","created":1776694000,"input_modalities":["text"],"output_modalities":["text"],"context_length":202752,"max_output_length":128000,"pricing":{"prompt":"0.00000121","completion":"0.0000042","input_cache_read":"0.0000006"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"zai-org/GLM-5.1","is_tee":true,"providers":["chutes"],"description":"GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...","metadata":{}},{"id":"anthropic/claude-opus-4.5","name":"Anthropic: Claude Opus 4.5","created":1764039380,"input_modalities":["file","image","text"],"output_modalities":["text"],"context_length":200000,"max_output_length":64000,"pricing":{"prompt":"0.000005","completion":"0.000025"},"supported_parameters":["include_reasoning","max_tokens","reasoning","response_format","stop","structured_outputs","temperature","tool_choice","tools","top_k","verbosity"],"supported_sampling_parameters":["max_tokens","stop","temperature","top_k"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"is_tee":false,"providers":["anthropic"],"description":"Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and reasoning benchmarks, and improved robustness to prompt injection. The model is designed to operate efficiently across varied effort levels, enabling developers to trade off speed, depth, and token usage depending on task requirements. It comes with a new parameter to control token efficiency, which can be accessed using the OpenRouter Verbosity parameter with low, medium, or high. Opus 4.5 supports advanced tool use, extended context management, and coordinated multi-agent setups, making it well-suited for autonomous research, debugging, multi-step planning, and spreadsheet/browser manipulation. It delivers substantial gains in structured reasoning, execution reliability, and alignment compared to prior Opus generations, while reducing token overhead and improving performance on long-running tasks.","metadata":{}},{"id":"deepseek/deepseek-v3.2","name":"DeepSeek: DeepSeek V3.2","created":1764804811,"input_modalities":["text"],"output_modalities":["text"],"context_length":163840,"max_output_length":64000,"pricing":{"prompt":"0.000001","completion":"0.000001","input_cache_read":"0.0000005"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"hugging_face_id":"deepseek-ai/DeepSeek-V3.2","is_tee":true,"providers":["chutes"],"description":"DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.","metadata":{}},{"id":"qwen/qwen3-vl-30b-a3b-instruct","name":"Qwen: Qwen3 VL 30B A3B Instruct","created":1764320402,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":128000,"max_output_length":32768,"pricing":{"prompt":"0.0000002","completion":"0.0000007"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3-VL-30B-A3B-Instruct","is_tee":true,"providers":["near-ai"],"description":"Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception of real-world/synthetic categories, 2D/3D spatial grounding, and long-form visual comprehension, achieving competitive multimodal benchmark results. For agentic use, it handles multi-image multi-turn instructions, video timeline alignments, GUI automation, and visual coding from sketches to debugged UI. Text performance matches flagship Qwen3 models, suiting document AI, OCR, UI assistance, spatial tasks, and agent research.","metadata":{}},{"id":"qwen/qwen-2.5-7b-instruct","name":"Qwen2.5 7B Instruct","created":1759532719,"input_modalities":["text"],"output_modalities":["text"],"context_length":32768,"max_output_length":32768,"pricing":{"prompt":"0.0000001","completion":"0.0000002"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","tools"],"hugging_face_id":"Qwen/Qwen2.5-7B-Instruct","is_tee":true,"providers":["phala"],"description":"Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2:\n\n- Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains.\n\n- Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and condition-setting for chatbots.\n\n- Long-context Support up to 128K tokens and can generate up to 8K tokens.\n\n- Multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more.\n\nUsage of this model is subject to [Tongyi Qianwen LICENSE AGREEMENT](https://huggingface.co/Qwen/Qwen1.5-110B-Chat/blob/main/LICENSE).","metadata":{}},{"id":"moonshotai/kimi-k3","name":"MoonshotAI: Kimi K3","created":1785294047,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":1048576,"max_output_length":1048576,"pricing":{"prompt":"0.000003","completion":"0.000015","input_cache_read":"0.0000003"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"moonshotai/Kimi-K3","is_tee":true,"providers":["chutes"],"description":"Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows.","metadata":{}},{"id":"xiaomi/mimo-v2.5-pro","name":"Xiaomi: MiMo-V2.5-Pro","created":1785329883,"input_modalities":["text"],"output_modalities":["text"],"context_length":1050000,"max_output_length":1050000,"pricing":{"prompt":"0.000000522","completion":"0.000001044","input_cache_read":"0.0000000048","input_cache_write":"0"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"quantization":"bf16","is_tee":false,"providers":["redpill"],"description":"MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....","metadata":{}},{"id":"xiaomi/mimo-v2.5","name":"Xiaomi: MiMo-V2.5","created":1785329882,"input_modalities":["text","audio","image","video"],"output_modalities":["text"],"context_length":1050000,"max_output_length":1050000,"pricing":{"prompt":"0.000000168","completion":"0.000000336","input_cache_read":"0.0000000036","input_cache_write":"0"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"quantization":"fp8","is_tee":false,"providers":["redpill"],"description":"MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...","metadata":{}}]}