Skip to content

Local LLM 与硬件加速 ​

#LLM · #本地推理 · #硬件加速 · #量化 · #AppleSilicon · #显存 · #吞吐量

当你需要用 HTML+JS 写一个本地轻量级 Web 推理界面时,底层需要对硬件有深刻理解。本专题聚焦本地运行大语言模型的硬件支撑、性能调优与量化选择。


统一内存(Unified Memory)架构 ​

传统架构 vs 统一内存 ​

mermaid
graph TD
    subgraph "传统 PC 架构 (x86 + 独显)"
        CPU["CPU<br/>i9-14900K"] --- DDR["DDR5 系统内存<br/>64-128GB"]
        CPU --- PCIE["PCIe 4.0 x16<br/>~32 GB/s"]
        PCIE --- GPU["NVIDIA RTX 4090"]
        GPU --- VRAM["GDDR6X 显存<br/>24GB"]
    end

    subgraph "Apple Silicon (M1/M2/M3 Max/Ultra)"
        SOC["Apple M3 Max SoC<br/>统一内存控制器"]
        SOC --- UM1["CPU 可访问"]
        SOC --- UM2["GPU 可访问"]
        SOC --- UM3["ANE (神经网络引擎) 可访问"]
        UM1 --- MEM["统一内存 LPDDR5<br/>最高 128GB<br/>带宽 400 GB/s"]
        UM2 --- MEM
        UM3 --- MEM
    end

    style MEM fill:#e07b39,color:#fff
    style VRAM fill:#e74c3c,color:#fff

关键区别一览 ​

特性传统 x86 + 独显Apple Silicon 统一内存
CPU ↔ GPU 数据传输通过 PCIe 总线,需显式拷贝零拷贝,共享同一物理内存
最大可用"显存"受 VRAM 物理限制(24/48/80GB)受统一内存限制(最高 128/192GB)
带宽PCIe 4.0: ~32 GB/s;显存: ~1 TB/sM3 Max: 400 GB/s; M2 Ultra: 800 GB/s
GPU 编程模型CUDA (NVIDIA) / ROCm (AMD)Metal (Apple) / MPS (PyTorch)
大模型友好度VRAM 瓶颈,需要多卡内存即显存,可直接加载大模型

Apple Silicon 内存分配策略 ​

python
"""
Apple Silicon 上 LLM 推理时内存管理的关键考虑
"""

import torch
import platform

def check_memory_availability():
    """检查 Apple Silicon 上可分配给 GPU 的内存"""
    if platform.system() != "Darwin":
        return

    # MPS (Metal Performance Shaders) 后端检测
    if torch.backends.mps.is_available():
        # 获取系统总内存(不是显存,因为统一内存共享)
        import subprocess
        result = subprocess.run(
            ["sysctl", "-n", "hw.memsize"],
            capture_output=True, text=True
        )
        total_bytes = int(result.stdout.strip())
        total_gb = total_bytes / (1024**3)

        # 获取当前可用内存
        vm_result = subprocess.run(
            ["vm_stat"],
            capture_output=True, text=True
        )

        print(f"系统总统一内存: {total_gb:.1f} GB")
        print(f"理论上大模型可用: {total_gb * 0.7:.1f} GB (预留30%给系统)")

        # 实际分配建议:保守使用不超过 70% 总内存
        max_model_memory = total_bytes * 0.7
        return max_model_memory

# 根据模型大小估算所需内存
def estimate_model_memory(params_billions: float, quantization: str = "fp16"):
    """估算模型加载到内存所需的大小

    Args:
        params_billions: 参数量(十亿)
        quantization: 量化方式 fp32 / fp16 / q8 / q4

    Returns:
        所需内存(GB)
    """
    bytes_per_param = {
        "fp32": 4,
        "fp16": 2,
        "q8":   1,    # 8-bit 量化 ≈ 1 字节每参数
        "q4":   0.5,  # 4-bit 量化 ≈ 0.5 字节每参数
    }

    base_memory = params_billions * bytes_per_param.get(quantization, 2)

    # 加上 KV Cache、激活值等额外开销 (约 20%)
    total_memory = base_memory * 1.2

    return total_memory

# 示例
for q in ["fp16", "q8", "q4"]:
    mem = estimate_model_memory(70, q)  # LLaMA 4 Scout
    print(f"LLaMA 4 Scout@{q}: 需要约 {mem:.1f} GB")

避免 OOM(内存溢出)的策略 ​

mermaid
graph TD
    A["加载 LLM 开始"] --> B{"参数检查<br/>模型大小 < 可用内存?"}
    B -->|否| C["❌ 无法加载<br/>尝试更小模型或更高压缩量化"]
    B -->|是| D{"预留 30% 系统内存<br/>避免系统 OOM"}
    D -->|不足| C
    D -->|充足| E["加载模型"]
    E --> F["设置内存上限<br/>PYTORCH_MPS_HIGH_WATERMARK_RATIO"]
    F --> G["推理时监控<br/>memory_pressure / vm_stat"]

    style C fill:#e74c3c,color:#fff
    style G fill:#2ecc71,color:#fff
bash
# Apple Silicon 上限制 PyTorch MPS 内存使用的环境变量
export PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.5   # 最多使用 50% 统一内存
export PYTORCH_MPS_LOW_WATERMARK_RATIO=0.3    # 低水位线 30%

# 实时监控内存压力
watch -n 1 'memory_pressure | grep "System-wide memory free percentage"'

陷阱:在 Apple Silicon 上,torch.cuda.max_memory_allocated() 不可用。需使用 torch.mps.current_allocated_memory() 和系统级 memory_pressure 命令来监控。


吞吐量基准测试与调优 ​

LLM 推理的两个阶段 ​

mermaid
sequenceDiagram
    participant User as 用户
    participant Engine as 推理引擎 (llama.cpp / vLLM)

    Note over User,Engine: 🔥 Prefill 阶段(一次完成)
    User->>Engine: 发送完整 Prompt(如整篇 HTML 文档)
    Engine->>Engine: 并行计算所有 Token 的 KV Cache
    Note over Engine: 瓶颈:计算密集(Compute-bound)<br/>受 GPU FLOPS 限制

    Note over User,Engine: 📝 Decode 阶段(逐 Token 循环)
    loop 每个生成 Token
        Engine->>Engine: 计算下一个 Token 概率
        Engine->>User: 输出 1 个 Token
        Note over Engine: 瓶颈:内存带宽(Memory-bound)<br/>受显存带宽限制
    end
阶段英文名特点瓶颈受什么限制
预填充Prefill一次处理全部 Prompt计算密集 (Compute-bound)GPU FLOPS
解码Decode逐 Token 自回归生成内存带宽 (Memory-bound)显存/统一内存带宽

Prefill 阶段优化 ​

当用户粘贴一篇 20KB 的 HTML 文档作为 Prompt 时,Prefill 阶段一次性处理所有 Token。

python
"""
Prefill 阶段性能分析
"""
import time
import torch

def benchmark_prefill(model, tokenizer, prompt: str, num_runs: int = 5):
    """测试 Prompt 处理的吞吐量"""
    inputs = tokenizer(prompt, return_tensors="pt")

    # 预热
    with torch.no_grad():
        _ = model(**inputs)

    # 正式测试
    times = []
    for _ in range(num_runs):
        start = time.perf_counter()
        with torch.no_grad():
            outputs = model(**inputs)
        # Apple Silicon MPS 需要显式同步
        if torch.backends.mps.is_available():
            torch.mps.synchronize()
        times.append(time.perf_counter() - start)

    avg_time = sum(times) / len(times)
    num_tokens = inputs["input_ids"].shape[1]
    tokens_per_second = num_tokens / avg_time

    print(f"Prompt 长度: {num_tokens} tokens")
    print(f"Prefill 耗时: {avg_time*1000:.0f} ms")
    print(f"Prefill 吞吐: {tokens_per_second:.0f} tokens/s")
    return tokens_per_second

Decode 阶段的 KV Cache 优化 ​

KV Cache 是 LLM 推理最大的内存消耗者之一。随着对话增长,KV Cache 不断膨胀:

KV Cache 大小 = 2 × 层数 × 头数 × 每头维度 × 序列长度 × 量化字节数

示例(Llama 2 7B,FP16 KV Cache):

  • 层数: 32, 头数: 32, 每头维度: 128
  • 每 1000 Token 的 KV Cache 大小:2 × 32 × 32 × 128 × 1000 × 2 bytes ≈ 512 MB
python
def calculate_kv_cache_size(
    num_layers: int,
    num_heads: int,
    head_dim: int,
    seq_len: int,
    kv_bits: int = 16
) -> float:
    """计算 KV Cache 占用内存

    Args:
        num_layers: Transformer 层数
        num_heads: 每层注意力头数(注意 GQA 情况下是 KV 头数)
        head_dim: 每头维度
        seq_len: 当前序列长度
        kv_bits: KV Cache 精度 (16/8/4)
    """
    bytes_per_element = kv_bits / 8
    # ×2 是因为 K 和 V 各一份
    size_bytes = 2 * num_layers * num_heads * head_dim * seq_len * bytes_per_element
    return size_bytes / (1024**3)  # 转为 GB


# 对比不同模型的 KV Cache 增长
models = [
    ("Llama 2 7B", 32, 32, 128),       # MHA
    ("LLaMA 4 Scout", 32, 8, 128),        # GQA (KV 头=8)
    ("LLaMA 4 Scout", 80, 8, 128),       # GQA
]

for name, layers, kv_heads, dim in models:
    for seq_k in [1, 4, 8, 32, 100]:
        cache_gb = calculate_kv_cache_size(layers, kv_heads, dim, seq_k * 1024)
        print(f"{name} @ {seq_k}K tokens: KV Cache = {cache_gb:.2f} GB")

长对话下防止吞吐量断崖下跌 ​

mermaid
graph TD
    A["对话进行中<br/>KV Cache 持续增长"] --> B{"KV Cache 是否<br/>超过预算? (如 4GB)"}
    B -->|否| C["✅ 继续正常推理"]
    B -->|是| D["触发 KV Cache 管理策略"]
    D --> E["策略1: Rolling Cache<br/>滑窗式丢弃最早Token"]
    D --> F["策略2: 压缩/量化<br/>将 FP16 KV 压缩为 Q4/Q8"]
    D --> G["策略3: 摘要归并<br/>将历史对话压缩为摘要"]
    E --> C
    F --> C
    G --> C

    style B fill:#f39c12,color:#fff
    style D fill:#e74c3c,color:#fff
python
"""
KV Cache 管理的关键参数配置(llama.cpp 视角)
"""

# === LLM 推理服务的推荐参数 ===

CONFIG = {
    # 上下文窗口配置
    "ctx_size": 32768,         # 最大上下文长度 (32K)
    "batch_size": 512,         # Prefill 批处理大小

    # KV Cache 相关
    "cache_type_k": "f16",     # K 精度: f16 | q8_0 | q4_0
    "cache_type_v": "f16",     # V 精度: f16 | q8_0 | q4_0
    # 当 KV Cache 使用 Q8 时,显存节省 50%,质量几乎无损失

    # 防止 OOM 的内存限制
    "memory_f16": True,        # f16 KV Cache 需要更大内存
    "use_mlock": True,         # 锁定内存,防止 swap 影响性能

    # 长对话参数
    "rope_freq_base": 1000000, # NTK-aware 扩展,提升长上下文效果
    "rope_freq_scale": 1.0,

    # Flash Attention(减少 KV Cache 中间内存)
    "flash_attn": True,        # 推荐开启,降低显存峰值
}

# 动态批处理建议
DYNAMIC_BATCH = """
长对话策略:
  短对话 (<2K tokens):  batch=512, KV FP16
  中对话 (2K-8K):       batch=256, KV Q8
  长对话 (>8K):          batch=128, KV Q8, 同时启用 Rolling Cache
"""

量化精度的工程取舍 ​

量化方案对比 ​

量化是指将模型参数从高精度(FP16/FP32)压缩到低精度(INT8/INT4),以减少内存占用和加速推理。

mermaid
graph LR
    FP32["FP32<br/>32 位浮点<br/>~4 bytes/param"] -->|"量化为"| FP16["FP16<br/>16 位浮点<br/>~2 bytes/param"]
    FP16 -->|"量化为"| Q8["Q8_0<br/>8 位整数<br/>~1 byte/param"]
    Q8 -->|"量化为"| Q4["Q4_K_M<br/>4 位整数<br/>~0.5 bytes/param"]
    Q4 -->|"量化为"| Q2["Q2_K<br/>2 位整数<br/>~0.25 bytes/param"]

    style FP16 fill:#2ecc71,color:#fff
    style Q4 fill:#f39c12,color:#fff
    style Q2 fill:#e74c3c,color:#fff

主流量化格式对比 ​

格式每参数字节7B 模型大小70B 模型大小质量损失适用场景
FP162 bytes~14 GB~140 GB无(baseline)GPU 推理 / 训练
Q8_0~1 byte~7 GB~70 GB几乎无感高精度推理
Q6_K~0.75 byte~5.5 GB~55 GB轻微需要高精度但内存紧张
Q5_K_M~0.65 byte~5 GB~48 GB可接受平衡选择
Q4_K_M~0.55 byte~4.2 GB~42 GB轻微退化最推荐的本地部署格式
Q4_0~0.5 byte~3.9 GB~40 GB有一定退化旧格式,不推荐新项目
Q3_K_M~0.4 byte~3 GB~30 GB明显退化极端内存受限
Q2_K~0.25 byte~2.5 GB~25 GB严重退化仅用于实验/边缘设备

Q4_K_M vs Q8_0 vs FP16:实测对比 ​

python
"""
量化精度对生成质量的影响 —— 实测对比框架

注意:这只是测试框架,实际数据需要在相同硬件和模型上跑 benchmark
"""

import time
from dataclasses import dataclass
from typing import List, Dict, Optional

@dataclass
class QuantBenchmark:
    """量化精度基准测试"""
    model_name: str           # 模型名称
    quantization: str         # 量化格式
    model_size_gb: float      # 模型文件大小
    memory_usage_gb: float    # 推理时实际内存占用
    prefill_speed: float      # tokens/s (Prefill 吞吐)
    decode_speed: float       # tokens/s (Decode 速度)
    perplexity: float         # 困惑度 (WikiText-2)

# 基于 llama.cpp 社区的实测数据(Llama 2 7B 典型值)
BENCHMARKS: List[QuantBenchmark] = [
    # 格式: (模型, 量化, 文件大小, 内存占用, Prefill速度, Decode速度, 困惑度)
    QuantBenchmark("Llama-2-7B", "FP16", 13.5, 16.2, 3200, 48,  5.47),
    QuantBenchmark("Llama-2-7B", "Q8_0", 7.2,  9.5,  3800, 62,  5.48),
    QuantBenchmark("Llama-2-7B", "Q6_K", 5.8,  8.0,  4100, 67,  5.50),
    QuantBenchmark("Llama-2-7B", "Q5_K_M", 5.0, 7.2,  4300, 71,  5.55),
    QuantBenchmark("Llama-2-7B", "Q4_K_M", 4.1, 6.3,  4800, 78,  5.65),
    QuantBenchmark("Llama-2-7B", "Q4_0",  3.8,  6.0,  4900, 82,  5.85),
    QuantBenchmark("Llama-2-7B", "Q2_K",  2.8,  4.8,  5500, 95,  6.80),
]

def compare_quantizations(benchmarks: List[QuantBenchmark]):
    """打印量化对比表"""
    header = f"{'量化':<8} {'大小(GB)':<10} {'内存(GB)':<10} {'Prefill t/s':<13} {'Decode t/s':<13} {'困惑度':<8}"
    print(header)
    print("-" * len(header))

    for b in benchmarks:
        print(f"{b.quantization:<8} {b.model_size_gb:<10.1f} {b.memory_usage_gb:<10.1f} "
              f"{b.prefill_speed:<13.0f} {b.decode_speed:<13.0f} {b.perplexity:<8.2f}")

# 运行对比
# compare_quantizations(BENCHMARKS)

工程选择建议 ​

代码生成场景 (高精度要求):
  优先: Q8_0 / Q6_K
  理由: 代码语法对精度敏感,Q4 容易出现括号不匹配、语法错误

通用对话场景:
  优先: Q4_K_M
  理由: 性价比之王,质量损失可接受,内存占用减半

逻辑推理 / 数学:
  优先: Q8_0 或 FP16
  理由: 推理链不对细节极其敏感,低精度下链式推理容易偏离

翻译场景:
  优先: Q5_K_M / Q4_K_M
  理由: 翻译对精度要求中等,量化影响较小

边缘设备 (树莓派/手机):
  优先: Q2_K 或 Q3_K_M
  理由: 聊胜于无,至少能跑起来

量化格式命名规则 ​

Q {位数} _ {K-quants} _ {混合精度}

Q4_K_M:
  Q4  = 4 位量化
  K   = K-quants 新算法(比旧 Q4_0 好很多)
  M   = Medium(中等大小,还有 S=Small, L=Large)

关键原则:
- 「_0」后缀(如 Q4_0)= 旧格式 → 不推荐
- 「_K_M」= 推荐日常使用
- 「_K_S」= 更小但质量稍降
- 「_K_L」= 质量更好但更大
- 「_0」vs「_1」:_0 是 0.5 bit 超分块(如 Q4_0 每参数正好 4.5 bit),
  _1 是 1 bit 超分块

本地 Web 推理界面的硬件选型指南 ​

mermaid
graph TD
    START["选择本地 LLM 部署方案"] --> Q1{"你的硬件?"}

    Q1 -->|"Apple Silicon<br/>M1/M2/M3"| A1["✅ 统一内存架构<br/>推荐方案"]
    Q1 -->|"NVIDIA GPU<br/>>16GB VRAM"| A2["✅ CUDA 生态<br/>最成熟"]
    Q1 -->|"Intel/AMD CPU<br/>无独显"| A3["⚠️ CPU 推理<br/>速度较慢但可行"]

    A1 --> B1["llama.cpp + Metal<br/>Q4_K_M 模型<br/>HTML+JS 前端"]
    A2 --> B2["llama.cpp + CUDA<br/>或 vLLM / TGI<br/>Q4_K_M ~ FP16"]
    A3 --> B3["llama.cpp CPU only<br/>Q4_K_M / Q4_0<br/>降低预期速度"]

    B1 --> WEB["🎯 本地 Web 推理界面<br/>HTML+JS 调用本地 API"]
    B2 --> WEB
    B3 --> WEB

    style A1 fill:#2ecc71,color:#fff
    style A2 fill:#2ecc71,color:#fff
    style A3 fill:#f39c12,color:#fff

推荐的本地部署技术栈 ​

场景推理引擎量化格式前端适合模型
MacBook Pro M3 Max (36GB)llama.cpp (Metal)Q4_K_MHTML+JS → 本地 APILLaMA 4 Scout 17Bx16E / Qwen4-72B
RTX 4090 (24GB)llama.cpp (CUDA) / vLLMQ4_K_M / FP16HTML+JS → 本地 APILLaMA 4 Scout / Qwen4-14B
MacBook Air M1 (8GB)llama.cpp (Metal)Q4_0HTML+JS → 本地 APIPhi-4 Mini / Gemma 3 4B
无 GPU (32GB DDR5)llama.cpp (CPU)Q4_K_MHTML+JS → 本地 APIQwen4-7B (慢, ~5 tok/s)

核心结论:Apple Silicon 的统一内存是最佳的本地 LLM 推理平台之一,因为内存即显存,可以以较低成本运行 70B+ 参数的大模型。


实战:llama.cpp 与 Ollama 部署 ​

llama.cpp — 最通用的 C++ 推理引擎 ​

bash
# ===== 1. 编译 llama.cpp =====
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp

# macOS(Metal 加速)
make -j LLAMA_METAL=1

# Linux(CUDA 加速)
make -j LLAMA_CUDA=1

# CPU only
make -j

# ===== 2. 下载量化模型(以 Qwen4-7B Q4_K_M 为例)=====
# 方法A: 从 HuggingFace 下载
huggingface-cli download Qwen/Qwen4-7B-Instruct-GGUF qwen4-7b-instruct-q4_k_m.gguf \
  --local-dir ./models

# 方法B: 自行转换原始模型为 GGUF(需要 convert.py)
python convert_hf_to_gguf.py /path/to/Qwen4-7B --outtype q4_k_m

# ===== 3. 启动推理服务器 =====
./llama-server \
  -m ./models/qwen4-7b-instruct-q4_k_m.gguf \
  -c 4096 \              # 上下文长度
  -ngl 99 \              # GPU 层数(Metal/CUDA 时建议全部层放 GPU)
  -b 512 \               # Prefill batch size
  --host 0.0.0.0 \       # 监听所有网卡
  --port 8080 \          # 端口
  --api-key sk-local     # API 密钥(可选)

# ===== 4. 测试 API =====
curl http://localhost:8080/v1/chat/completions \
  -H "Authorization: Bearer sk-local" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-3.5-turbo",
    "messages": [{"role": "user", "content": "什么是 GMP 调度模型?"}],
    "temperature": 0.7,
    "max_tokens": 512
  }'

# ===== 5. 性能测试 =====
# 生成 512 token 的基准测试
time curl -s http://localhost:8080/v1/completions \
  -H "Authorization: Bearer sk-local" \
  -d '{"prompt": "请用500字介绍Go语言", "max_tokens": 512}' \
  | jq -r '.content'
# 记录 tokens/s: 观察服务器日志中的 eval time

Ollama — 开箱即用的 LLM 运行时 ​

bash
# ===== 1. 安装 Ollama =====
# macOS: https://ollama.com/download
# Linux:
curl -fsSL https://ollama.com/install.sh | sh

# ===== 2. 拉取模型 =====
ollama pull qwen4:7b             # 7B 模型(~4.7GB)
ollama pull qwen4:14b            # 14B 模型(~8.5GB)
ollama pull qwen4:32b            # 32B 模型(~19GB)
ollama pull llama4-scout:17bx16e # LLaMA 4 Scout MoE
#   Radix: 100 token Prefill(节省 83%)
ollama pull deepseek-r2:8b        # DeepSeek-R2 推理模型

# 查看已下载的模型
ollama list
# NAME                ID              SIZE      MODIFIED
# qwen4:7b            845dbda0ea48    4.7 GB    2 days ago
# llama4-scout        a80c4f17acd5    5.2 GB    5 days ago

# ===== 3. 启动 Ollama 服务(默认自动启动) =====
ollama serve

# ===== 4. 命令行交互 =====
ollama run qwen4:7b
# >>> 什么是 GMP 调度模型?
# ...(输出回答)
# >>> /bye

# ===== 5. API 调用(兼容 OpenAI 格式) =====
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen4:7b",
    "messages": [{"role": "user", "content": "解释GMP调度模型"}],
    "temperature": 0.1,
    "stream": false
  }'

# ===== 6. 自定义 Modelfile(微调系统提示词) =====
cat > Modelfile << 'EOF'
FROM qwen4:7b

# 设定温度和上下文窗口
PARAMETER temperature 0.1
PARAMETER num_ctx 4096

# 系统提示词
SYSTEM """你是一个编程专家助手,精通 Go、Python、Rust、JavaScript。
回答问题时给出准确、可运行的代码示例。
对于不确定的问题,明确说明不确定。"""

# 停止序列
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
EOF

# 创建自定义模型
ollama create my-coding-assistant -f Modelfile
ollama run my-coding-assistant

# ===== 7. 性能调优 =====
# 查看 GPU 使用情况
ollama ps
# NAME                ID              SIZE      PROCESSOR    UNTIL
# qwen4:7b          845dbda0ea48    5.6 GB    100% GPU     4 minutes from now

# 设置并发和队列
export OLLAMA_NUM_PARALLEL=4      # 最大并发请求
export OLLAMA_MAX_QUEUE=512       # 请求队列长度
export OLLAMA_KEEP_ALIVE=5m       # 模型在内存中的保持时间

llama.cpp vs Ollama 对比 ​

维度llama.cppOllama
安装难度需编译源码curl | sh 一键安装
模型管理手动下载 ggufollama pull 自动管理
性能完全可控,极致优化基于 llama.cpp,自动调优
API 兼容OpenAI 兼容(llama-server)OpenAI 兼容
生产环境✅ 推荐(精细控制)✅ 推荐(运维简单)
开发调试需要更多配置✅ 开箱即用
自定义模型完全自由Modelfile 配置文件

用 vLLM 部署生产级推理服务 ​

bash
# vLLM 适合有 NVIDIA GPU 的生产环境
pip install vllm

# 启动服务
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen4-7B-Instruct \
  --dtype float16 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.9 \
  --tensor-parallel-size 1 \       # 单卡=1,双卡=2
  --host 0.0.0.0 \
  --port 8000

# 性能: 7B 模型在 RTX 4090 上可达 100+ tokens/s
# vLLM 特色功能:
# - PagedAttention: 高效管理 KV Cache
# - Continuous batching: 动态批处理,最大化吞吐
# - Tensor parallelism: 多卡自动切分

实战:搭建本地 LLM + Web UI 的完整流程 ​

bash
# === 方案一:Ollama + Open WebUI(最简方案) ===

# 1. 启动 Ollama
ollama pull qwen4:7b
ollama serve &

# 2. 启动 Open WebUI(Docker)
docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main

# 3. 访问 http://localhost:3000
# Web UI 自带对话历史、文档上传、模型切换


# === 方案二:Ollama + 自定义 FastAPI 后端 ===

# app.py
from fastapi import FastAPI
from openai import OpenAI

app = FastAPI()
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

@app.post("/chat")
async def chat(question: str):
    response = client.chat.completions.create(
        model="qwen4:7b",
        messages=[{"role": "user", "content": question}],
        temperature=0.1,
        max_tokens=1024,
    )
    return {"answer": response.choices[0].message.content}

# 启动: uvicorn app:app --port 8000
批注模式

💬 文章评论

暂无评论,来说点什么吧 👇

编程学习笔记