与 AI 模型进行多轮对话,支持流式输出。这是最常用的接口,OpenAI 完全兼容。
Chat 对话 是 AI 应用的核心接口:把你的消息发给模型,模型返回回复。支持多轮对话(传历史消息)、角色设定(system)、流式打字机效果。
有什么用:聊天机器人、AI 助手、客服、写作、翻译、代码生成等一切需要 AI 生成文本的场景。
| 项目 | 值 |
|---|---|
| 请求方法 | POST |
| 请求路径 | https://103.236.87.35/v1/chat/completions |
| 请求头 | Authorization: Bearer <你的密钥>Content-Type: application/json |
| 是否需要鉴权 | 必需 |
| 是否计费 | 按 Token 计费 |
| 参数名 | 类型 | 必填 | 默认值 | 说明 |
|---|---|---|---|---|
model | string | 是 | - | 模型名称,如 deepseek-v4-flash。见「模型列表」接口 |
messages | array | 是 | - | 对话消息数组,每项含 role 和 content |
messages[].role | string | 是 | - | 角色:system(系统设定)/ user(用户)/ assistant(助手) |
messages[].content | string | 是 | - | 消息内容文本 |
temperature | number | 否 | 1 | 随机性,0~2。越低越稳定,越高越有创造性 |
top_p | number | 否 | 1 | 核采样,与 temperature 二选一调整 |
max_tokens | integer | 否 | 模型默认 | 回复最大 token 数 |
stream | boolean | 否 | false | 是否流式返回(SSE),true 时逐块输出 |
stop | string/array | 否 | - | 遇到该字符串即停止生成 |
presence_penalty | number | 否 | 0 | 话题新鲜度惩罚,-2~2 |
frequency_penalty | number | 否 | 0 | 重复度惩罚,-2~2 |
user | string | 否 | - | 终端用户标识 |
curl https://103.236.87.35/v1/chat/completions \
-H "Authorization: Bearer <你的密钥>" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "system", "content": "你是一个乐于助人的助手"},
{"role": "user", "content": "你好,介绍一下你自己"}
],
"temperature": 0.7,
"max_tokens": 1024
}'
from openai import OpenAI
client = OpenAI(
api_key="<你的密钥>",
base_url="https://103.236.87.35/v1",
)
resp = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{"role": "system", "content": "你是一个乐于助人的助手"},
{"role": "user", "content": "你好,介绍一下你自己"},
],
)
print(resp.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "<你的密钥>",
baseURL: "https://103.236.87.35/v1",
});
const resp = await client.chat.completions.create({
model: "deepseek-v4-flash",
messages: [{ role: "user", content: "你好" }],
});
console.log(resp.choices[0].message.content);
{
"id": "chatcmpl-0217888408922",
"object": "chat.completion",
"created": 1788840893,
"model": "deepseek-v4-flash-ga-260731",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "你好!我是你的 AI 助手。"},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 86,
"completion_tokens": 35,
"total_tokens": 121
}
}
| 字段名 | 类型 | 说明 |
|---|---|---|
id | string | 本次请求唯一 ID |
object | string | 对象类型,固定 chat.completion |
created | integer | 创建时间戳(秒) |
model | string | 实际使用的模型 |
choices[].message.role | string | 回复角色,固定 assistant |
choices[].message.content | string | 模型回复的文本内容 |
choices[].finish_reason | string | 结束原因:stop 正常结束 / length 达到上限 / content_filter 内容过滤 |
usage.prompt_tokens | integer | 输入 token 数(计费依据) |
usage.completion_tokens | integer | 输出 token 数(计费依据) |
usage.total_tokens | integer | 总 token 数 |
设置 "stream": true,响应以 text/event-stream 逐块返回。网关会自动注入 stream_options.include_usage(流结束带计费数据),并剥离模型思考内容。
from openai import OpenAI
client = OpenAI(api_key="<你的密钥>", base_url="https://103.236.87.35/v1")
stream = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "讲个笑话"}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
{
"error": {
"message": "Invalid API key provided",
"type": "new_api_error",
"param": "",
"code": "invalid_api_key"
}
}
常见错误码:401 invalid_api_key 密钥错误 · 402 insufficient_balance 余额不足 · 403 model_not_allowed 模型无权 · 429 rate_limit_exceeded 调用超限。完整对照见文档首页 · 错误格式。