Skip to content
第 24 章 后端 ⏱ 9 分钟阅读

第 24 章:日志与监控 ​

学习目标 ​

  • 用 pino 替代 winston 做生产日志
  • 集成 Prometheus 暴露指标
  • 用 @nestjs/terminus 做健康检查
  • 接入 OpenTelemetry 做链路追踪

一、生产日志:pino(主流选择) ​

NestJS 默认 winston,但 pino 更快(10 倍),JSON 结构化输出,生产首选。

bash
pnpm add nestjs-pino pino-http
typescript
// app.module.ts
import { LoggerModule } from 'nestjs-pino';

@Module({
  imports: [
    LoggerModule.forRoot({
      pinoHttp: {
        level: process.env.NODE_ENV === 'production' ? 'info' : 'debug',
        transport: process.env.NODE_ENV === 'production'
          ? undefined                              // 生产:JSON 直发
          : { target: 'pino-pretty' },             // 开发:彩色
        redact: ['req.headers.authorization'],     // 隐藏敏感头
      },
    }),
  ],
})
export class AppModule {}
typescript
// 任意地方注入
import { PinoLogger } from 'nestjs-pino';

@Injectable()
export class UsersService {
  constructor(private readonly logger: PinoLogger) {}

  async create(dto: CreateUserDto) {
    this.logger.info('creating user', { username: dto.username });
    // 输出:
    // {"level":30,"time":1693645200000,"context":"UsersService","msg":"creating user","username":"tom"}
  }
}

⚠️ 坑 1:开发开 debug、生产开 info → 用环境变量切换。

二、请求日志 + traceId ​

bash
pnpm install pino-http
typescript
LoggerModule.forRoot({
  pinoHttp: {
    // 自动注入 traceId(每个请求唯一)
    genReqId: (req) => req.headers['x-request-id'] ?? randomUUID(),
    customLogLevel: (req, res, err) => {
      if (res.statusCode >= 500 || err) return 'error';
      if (res.statusCode >= 400) return 'warn';
      return 'info';
    },
  },
});

每次请求自动生成 traceId,跨服务追踪时传 x-request-id 头即可串联。

三、Prometheus 指标(事实标准) ​

bash
pnpm add @willsoto/nestjs-prometheus prom-client
typescript
// app.module.ts
import { PrometheusModule } from '@willsoto/nestjs-prometheus';

@Module({
  imports: [
    PrometheusModule.register({
      defaultMetrics: { enabled: true },           // CPU/内存/事件循环
      defaultLabels: { service: 'my-app' },
    }),
  ],
})
export class AppModule {}

访问 http://localhost:3000/metrics 自动输出 Prometheus 格式:

# HELP process_cpu_user_seconds_total Total user CPU time spent in seconds.
# TYPE process_cpu_user_seconds_total counter
process_cpu_user_seconds_total 0.012
...

四、自定义业务指标 ​

typescript
// orders.controller.ts
import { Counter } from 'prom-client';
import { InjectMetric } from '@willsoto/nestjs-prometheus';

@Controller('orders')
export class OrdersController {
  constructor(
    @InjectMetric('orders_total') private readonly ordersCounter: Counter<string>,
  ) {}

  @Post()
  async create() {
    // 业务逻辑...
    this.ordersCounter.inc({ status: 'success' });  // 计数 +1
    // 输出:orders_total{status="success"} 1
  }
}

生产必加的几个指标:

指标类型含义
http_requests_totalCounter接口累计调用次数
http_request_duration_secondsHistogram响应时间分布
orders_total{status="success|fail"}Counter订单成功/失败
cache_hits_total / cache_misses_totalCounter缓存命中率

五、Grafana 监控面板(Prometheus 配合) ​

yaml
# docker-compose.yml
services:
  prometheus:
    image: prom/prometheus
    ports: ['9090:9090']
    volumes: ['./prometheus.yml:/etc/prometheus/prometheus.yml']
  grafana:
    image: grafana/grafana
    ports: ['3001:3000']
yaml
# prometheus.yml
scrape_configs:
  - job_name: 'nestjs'
    static_configs:
      - targets: ['host.docker.internal:3000']

Grafana → Data Source → Prometheus → 配面板。

六、健康检查(必备) ​

bash
pnpm add @nestjs/terminus
typescript
@Controller('health')
export class HealthController {
  constructor(
    private readonly health: HealthCheckService,
    private readonly db: TypeOrmHealthIndicator,
    private readonly redis: MicroserviceHealthIndicator,
  ) {}

  @Get()
  check() {
    return this.health.check([
      () => this.db.pingCheck('database'),
      () => this.redis.pingCheck('redis', { transport: Transport.REDIS, options: { host: '...', port: 6379 } }),
    ]);
  }

  @Get('liveness')                                // K8s liveness
  liveness() {
    return this.health.check([]);                 // 只检查进程活着
  }

  @Get('readiness')                               // K8s readiness
  readiness() {
    return this.check();                          // 检查依赖(DB、Redis)
  }
}

⚠️ 坑 2:/health 没区分 liveness 和 readiness → K8s 重启时可能误杀。

七、链路追踪:OpenTelemetry(主流) ​

bash
pnpm add @opentelemetry/api @opentelemetry/sdk-node @opentelemetry/auto-instrumentations-node @opentelemetry/exporter-trace-otlp-http
typescript
// tracing.ts(必须最先 import)
import { NodeSDK } from '@opentelemetry/sdk-node';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';

const sdk = new NodeSDK({
  traceExporter: new OTLPTraceExporter({ url: 'http://jaeger:4318/v1/traces' }),
  instrumentations: [getNodeAutoInstrumentations()],
});

sdk.start();
typescript
// main.ts
import './tracing';                                // 第一行
// ...其他

接入后 HTTP/MySQL/Redis 调用都有 trace,统一发到 Jaeger / SkyWalking 控制台看。

八、错误监控:Sentry ​

bash
pnpm add @sentry/node
typescript
// main.ts
import * as Sentry from '@sentry/node';

Sentry.init({
  dsn: 'https://xxx@sentry.io/123',
  environment: process.env.NODE_ENV,
  tracesSampleRate: 0.1,                            // 10% 性能采样
});

const app = await NestFactory.create(AppModule);
Sentry.setupNestJsErrorHandler(app);               // NestJS 错误自动上报

线上代码报错 → 邮件/Slack 通知 → 关联到具体请求的 traceId。

九、本章小结 ​

组件生产选型
日志库pino(快 10 倍,JSON 结构化)
指标Prometheus + prom-client
监控面板Grafana 接 Prometheus
健康检查@nestjs/terminus + liveness/readiness
链路追踪OpenTelemetry(自动埋点)
APMJaeger / SkyWalking / Tempo
错误监控Sentry(邮件/Slack 报警)
不用winston(慢)/ 自定义 Logger(标准接口)

生产三件套:pino + Prometheus + Sentry,基本覆盖 80% 监控需求。

⚠️ 坑 3:不接 Sentry 等错误监控 → 线上报错只能等用户反馈,往往已经很严重。

本站基于 VitePress 构建 · 由 StackHub 团队维护