第 24 章:日志与监控
学习目标
- 用 pino 替代 winston 做生产日志
- 集成 Prometheus 暴露指标
- 用 @nestjs/terminus 做健康检查
- 接入 OpenTelemetry 做链路追踪
一、生产日志:pino(主流选择)
NestJS 默认 winston,但 pino 更快(10 倍),JSON 结构化输出,生产首选。
bash
pnpm add nestjs-pino pino-httptypescript
// app.module.ts
import { LoggerModule } from 'nestjs-pino';
@Module({
imports: [
LoggerModule.forRoot({
pinoHttp: {
level: process.env.NODE_ENV === 'production' ? 'info' : 'debug',
transport: process.env.NODE_ENV === 'production'
? undefined // 生产:JSON 直发
: { target: 'pino-pretty' }, // 开发:彩色
redact: ['req.headers.authorization'], // 隐藏敏感头
},
}),
],
})
export class AppModule {}typescript
// 任意地方注入
import { PinoLogger } from 'nestjs-pino';
@Injectable()
export class UsersService {
constructor(private readonly logger: PinoLogger) {}
async create(dto: CreateUserDto) {
this.logger.info('creating user', { username: dto.username });
// 输出:
// {"level":30,"time":1693645200000,"context":"UsersService","msg":"creating user","username":"tom"}
}
}⚠️ 坑 1:开发开
debug、生产开info→ 用环境变量切换。
二、请求日志 + traceId
bash
pnpm install pino-httptypescript
LoggerModule.forRoot({
pinoHttp: {
// 自动注入 traceId(每个请求唯一)
genReqId: (req) => req.headers['x-request-id'] ?? randomUUID(),
customLogLevel: (req, res, err) => {
if (res.statusCode >= 500 || err) return 'error';
if (res.statusCode >= 400) return 'warn';
return 'info';
},
},
});每次请求自动生成 traceId,跨服务追踪时传 x-request-id 头即可串联。
三、Prometheus 指标(事实标准)
bash
pnpm add @willsoto/nestjs-prometheus prom-clienttypescript
// app.module.ts
import { PrometheusModule } from '@willsoto/nestjs-prometheus';
@Module({
imports: [
PrometheusModule.register({
defaultMetrics: { enabled: true }, // CPU/内存/事件循环
defaultLabels: { service: 'my-app' },
}),
],
})
export class AppModule {}访问 http://localhost:3000/metrics 自动输出 Prometheus 格式:
# HELP process_cpu_user_seconds_total Total user CPU time spent in seconds.
# TYPE process_cpu_user_seconds_total counter
process_cpu_user_seconds_total 0.012
...四、自定义业务指标
typescript
// orders.controller.ts
import { Counter } from 'prom-client';
import { InjectMetric } from '@willsoto/nestjs-prometheus';
@Controller('orders')
export class OrdersController {
constructor(
@InjectMetric('orders_total') private readonly ordersCounter: Counter<string>,
) {}
@Post()
async create() {
// 业务逻辑...
this.ordersCounter.inc({ status: 'success' }); // 计数 +1
// 输出:orders_total{status="success"} 1
}
}生产必加的几个指标:
| 指标 | 类型 | 含义 |
|---|---|---|
http_requests_total | Counter | 接口累计调用次数 |
http_request_duration_seconds | Histogram | 响应时间分布 |
orders_total{status="success|fail"} | Counter | 订单成功/失败 |
cache_hits_total / cache_misses_total | Counter | 缓存命中率 |
五、Grafana 监控面板(Prometheus 配合)
yaml
# docker-compose.yml
services:
prometheus:
image: prom/prometheus
ports: ['9090:9090']
volumes: ['./prometheus.yml:/etc/prometheus/prometheus.yml']
grafana:
image: grafana/grafana
ports: ['3001:3000']yaml
# prometheus.yml
scrape_configs:
- job_name: 'nestjs'
static_configs:
- targets: ['host.docker.internal:3000']Grafana → Data Source → Prometheus → 配面板。
六、健康检查(必备)
bash
pnpm add @nestjs/terminustypescript
@Controller('health')
export class HealthController {
constructor(
private readonly health: HealthCheckService,
private readonly db: TypeOrmHealthIndicator,
private readonly redis: MicroserviceHealthIndicator,
) {}
@Get()
check() {
return this.health.check([
() => this.db.pingCheck('database'),
() => this.redis.pingCheck('redis', { transport: Transport.REDIS, options: { host: '...', port: 6379 } }),
]);
}
@Get('liveness') // K8s liveness
liveness() {
return this.health.check([]); // 只检查进程活着
}
@Get('readiness') // K8s readiness
readiness() {
return this.check(); // 检查依赖(DB、Redis)
}
}⚠️ 坑 2:
/health没区分 liveness 和 readiness → K8s 重启时可能误杀。
七、链路追踪:OpenTelemetry(主流)
bash
pnpm add @opentelemetry/api @opentelemetry/sdk-node @opentelemetry/auto-instrumentations-node @opentelemetry/exporter-trace-otlp-httptypescript
// tracing.ts(必须最先 import)
import { NodeSDK } from '@opentelemetry/sdk-node';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
const sdk = new NodeSDK({
traceExporter: new OTLPTraceExporter({ url: 'http://jaeger:4318/v1/traces' }),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();typescript
// main.ts
import './tracing'; // 第一行
// ...其他接入后 HTTP/MySQL/Redis 调用都有 trace,统一发到 Jaeger / SkyWalking 控制台看。
八、错误监控:Sentry
bash
pnpm add @sentry/nodetypescript
// main.ts
import * as Sentry from '@sentry/node';
Sentry.init({
dsn: 'https://xxx@sentry.io/123',
environment: process.env.NODE_ENV,
tracesSampleRate: 0.1, // 10% 性能采样
});
const app = await NestFactory.create(AppModule);
Sentry.setupNestJsErrorHandler(app); // NestJS 错误自动上报线上代码报错 → 邮件/Slack 通知 → 关联到具体请求的 traceId。
九、本章小结
| 组件 | 生产选型 |
|---|---|
| 日志库 | pino(快 10 倍,JSON 结构化) |
| 指标 | Prometheus + prom-client |
| 监控面板 | Grafana 接 Prometheus |
| 健康检查 | @nestjs/terminus + liveness/readiness |
| 链路追踪 | OpenTelemetry(自动埋点) |
| APM | Jaeger / SkyWalking / Tempo |
| 错误监控 | Sentry(邮件/Slack 报警) |
| 不用 | winston(慢)/ 自定义 Logger(标准接口) |
生产三件套:pino + Prometheus + Sentry,基本覆盖 80% 监控需求。
⚠️ 坑 3:不接 Sentry 等错误监控 → 线上报错只能等用户反馈,往往已经很严重。