"一图胜千言"。pandas 内置了基于 Matplotlib 的绘图接口(.plot),几行代码就能生成折线图、柱状图、散点图、箱线图、直方图等常用图表。本篇文章带你掌握 pandas 绘图的全部常用技能,并能与 Matplotlib 无缝配合做细节美化。目录
1. 绘图环境准备
2. Series / DataFrame 绘图总览
3. 折线图 line
4. 柱状图 bar / barh
5. 直方图 hist
6. 箱线图 box
7. 散点图 scatter
8. 饼图 pie
9. 面积图 area
10. 与 Matplotlib 配合美化
11. 子图布局
12. 实战:销售数据可视化报告
13. 常见坑与注意事项
14. 本章小结与练习
1. 绘图环境准备
# 安装(如未安装)
# pip install matplotlib
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# 在 Jupyter 中显示图片
%matplotlib inline
# 支持中文显示(Windows/常见环境)
plt.rcParams["font.sans-serif"] = ["SimHei", "Microsoft YaHei", "Arial Unicode MS"]
plt.rcParams["axes.unicode_minus"] = False # 正常显示负号
不设置中文字体的话,图表中文会显示成方框。这是最常见的坑。
2. Series / DataFrame 绘图总览
df.plot(kind="类型", 参数...),kind 决定图表类型:
| kind | 图表 | 适用场景 |
|------|------|----------|
| line(默认) | 折线图 | 趋势变化(时间序列) |
| bar | 柱状图 | 分类对比 |
| barh | 横向柱状图 | 分类对比(类别名长时) |
| hist | 直方图 | 数值分布 |
| box | 箱线图 | 分布与离群值 |
| scatter | 散点图 | 两变量关系 |
| pie | 饼图 | 占比 |
| area | 面积图 | 累积趋势 |
| hexbin | 六边形密度图 | 大数据散点密度 |
| kde | 核密度估计图 | 平滑分布 |
准备数据:
np.random.seed(42)
df = pd.DataFrame({
"日期": pd.date_range("2025-01-01", periods=30, freq="D"),
"销量": np.random.randint(100, 300, 30),
"广告费": np.random.randint(20, 80, 30),
})
sales_by_city = pd.Series(
[1200, 950, 800, 650, 500],
index=["北京", "上海", "广州", "深圳", "杭州"],
)
3. 折线图 line
# 最简单:Series 或带数值列的 DataFrame
df.set_index("日期")["销量"].plot()
plt.title("每日销量趋势")
plt.xlabel("日期")
plt.ylabel("销量")
plt.show()
# 多列折线(自动生成图例)
df.set_index("日期")[["销量", "广告费"]].plot()
plt.show()
# 定制
df.set_index("日期")["销量"].plot(
figsize=(10, 5), # 图大小
color="red", # 颜色
linestyle="--", # 线型
linewidth=2, # 线宽
marker="o", # 数据点标记
grid=True, # 网格
title="销量走势",
)
plt.show()
4. 柱状图 bar / barh
# 垂直柱状图
sales_by_city.plot(kind="bar", color="steelblue")
plt.title("各城市销量")
plt.ylabel("销量")
plt.xticks(rotation=45) # 标签旋转,避免重叠
plt.show()
# 横向柱状图(类别多时更清晰)
sales_by_city.plot(kind="barh", color="orange")
plt.xlabel("销量")
plt.show()
# DataFrame 多列分组柱状图
df_month = pd.DataFrame({
"一月": [500, 400],
"二月": [550, 420],
}, index=["华东", "华南"])
df_month.plot(kind="bar")
plt.show()
# 堆叠柱状图
df_month.plot(kind="bar", stacked=True)
plt.show()
rot=0可让 x 轴标签不旋转:df.plot(kind="bar", rot=0)。
5. 直方图 hist
# 单列直方图
df["销量"].plot(kind="hist", bins=15, edgecolor="white")
plt.xlabel("销量")
plt.show()
# 多列直方图对比(透明叠加)
df[["销量", "广告费"]].plot(kind="hist", alpha=0.6, bins=15)
plt.show()
# 密度图(平滑版)
df["销量"].plot(kind="kde")
plt.show()
# 累积直方图
df["销量"].plot(kind="hist", cumulative=True, bins=15)
plt.show()
6. 箱线图 box
# 一列箱线图
df["销量"].plot(kind="box")
plt.show()
# 多列箱线图对比(数据分布一目了然)
df[["销量", "广告费"]].plot(kind="box")
plt.show()
# 按分组画箱线图
import seaborn as sns
# 或 pandas:分组后画
df["城市"] = np.random.choice(["A市", "B市", "C市"], 30)
df.boxplot(column="销量", by="城市")
plt.show()
箱线图解读:
┌─────────┐ ┌─ 上边缘(Q3 + 1.5*IQR)
│ ┌───┐ │
│ │ │ │ ┌─ 上四分位数 Q3(75%)
│ ├─●─┤ │ ├─ 中位数(50%)
│ │ │ │ └─ 下四分位数 Q1(25%)
│ └───┘ │
└─────────┘ └─ 下边缘(Q1 - 1.5*IQR)
○ ○ ○ 离群值
7. 散点图 scatter
# 两列散点
df.plot(kind="scatter", x="广告费", y="销量")
plt.show()
# 带颜色映射(第三维:点的大小/颜色)
df.plot(
kind="scatter",
x="广告费", y="销量",
s=df["销量"] / 10, # 点大小
c="广告费", # 颜色渐变
colormap="viridis",
)
plt.show()
# 带回归趋势线(需 seaborn)
import seaborn as sns
sns.regplot(data=df, x="广告费", y="销量")
plt.show()
8. 饼图 pie
sales_by_city.plot(kind="pie", autopct="%.1f%%", figsize=(6, 6))
plt.ylabel("") # 去掉默认 y 轴标签
plt.title("城市销售占比")
plt.show()
# 自定义颜色与起始角度
sales_by_city.plot(
kind="pie",
autopct="%.1f%%",
startangle=90, # 起始角度
colors=["#ff9999", "#66b3ff", "#99ff99", "#ffcc99", "#c2c2f0"],
)
plt.show()
饼图适合 5 个以内分类;分类多时用横向柱状图更清晰。
9. 面积图 area
# 面积图(累计趋势)
df.set_index("日期")[["销量"]].plot(kind="area")
plt.show()
# 多列堆叠面积图
df.set_index("日期")[["销量", "广告费"]].plot(kind="area", stacked=False, alpha=0.5)
plt.show()
10. 与 Matplotlib 配合美化
pandas 的 .plot() 返回的是 Matplotlib 的 Axes 对象,可以直接用 Matplotlib 的 API 继续美化:
ax = df.set_index("日期")["销量"].plot(
figsize=(10, 5),
color="#2E86AB",
linewidth=2,
)
# 标题、标签、图例
ax.set_title("2025 年 1 月每日销量", fontsize=16, pad=15)
ax.set_xlabel("日期", fontsize=12)
ax.set_ylabel("销量(件)", fontsize=12)
# 网格、边框
ax.grid(True, linestyle="--", alpha=0.6)
ax.spines["top"].set_visible(False)
ax.spines["right"].set_visible(False)
# 标注最大值
max_idx = df["销量"].idxmax()
ax.annotate(
f"最高 {df.loc[max_idx, '销量']}",
xy=(max_idx, df.loc[max_idx, "销量"]),
xytext=(max_idx, df.loc[max_idx, "销量"] + 20),
arrowprops=dict(arrowstyle="->"),
)
# 填充区间
ax.fill_between(df.index, df["销量"], alpha=0.15)
# 图例与保存
ax.legend(["销量"])
plt.tight_layout()
plt.savefig("销量图.png", dpi=150) # 保存高清图
plt.show()
11. 子图布局
# 多个子图
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
# 每个 axes 一个图
df["销量"].plot(kind="line", ax=axes[0, 0], title="折线")
df["销量"].plot(kind="hist", ax=axes[0, 1], title="直方图", bins=15)
df["销量"].plot(kind="box", ax=axes[1, 0], title="箱线图")
sales_by_city.plot(kind="bar", ax=axes[1, 1], title="柱状图")
plt.tight_layout()
plt.show()
# 一行多列
fig, axes = plt.subplots(1, 3, figsize=(15, 4))
df.plot(kind="line", x="日期", y="销量", ax=axes[0])
df.plot(kind="scatter", x="广告费", y="销量", ax=axes[1])
df.plot(kind="kde", y="销量", ax=axes[2])
plt.tight_layout()
plt.show()
ax= 参数可以把图"画进"指定的子图,是实现复杂排版的关键。12. 实战:销售数据可视化报告
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
plt.rcParams["font.sans-serif"] = ["SimHei", "Microsoft YaHei"]
plt.rcParams["axes.unicode_minus"] = False
# 模拟数据
np.random.seed(7)
n = 120
sales = pd.DataFrame({
"日期": pd.date_range("2025-01-01", periods=n, freq="D"),
"区域": np.random.choice(["华东", "华南", "华北", "西南"], n),
"销量": np.random.randint(50, 300, n).astype(float),
"成本": np.random.randint(30, 200, n).astype(float),
})
sales["利润"] = sales["销量"] - sales["成本"]
sales = sales.set_index("日期")
# 1. 总体趋势(7 日滚动)
fig, axes = plt.subplots(2, 2, figsize=(14, 9))
sales["销量"].plot(ax=axes[0, 0], title="每日销量")
sales["销量"].rolling(7).mean().plot(ax=axes[0, 0], label="7日均线", color="orange")
axes[0, 0].legend()
# 2. 区域对比
sales.groupby("区域")["销量"].sum().sort_values().plot(
kind="barh", ax=axes[0, 1], title="区域销量合计", color="steelblue"
)
# 3. 销量分布
sales["销量"].plot(kind="hist", bins=20, ax=axes[1, 0], title="销量分布", edgecolor="white")
# 4. 利润 vs 成本 散点
sales.plot(kind="scatter", x="成本", y="利润", ax=axes[1, 1], title="成本与利润关系")
plt.tight_layout()
plt.show()
# 5. 月度汇总柱状图
monthly = sales.resample("M")[["销量", "利润"]].sum()
monthly.plot(kind="bar", title="月度销量与利润")
plt.xticks(rotation=0)
plt.show()
13. 常见坑与注意事项
| 坑 | 现象 | 解决办法 |
|----|------|----------|
| 中文乱码 | 方框/乱码 | 设置 font.sans-serif,含 SimHei/微软雅黑 |
| 负号显示为方块 | - 变方块 | axes.unicode_minus = False |
| Jupyter 不显示图 | 无输出 | %matplotlib inline 或 plt.show() |
| 多图重叠 | 标签挤在一起 | plt.tight_layout()、figsize、rot |
| 索引是日期时 x 轴乱 | 默认按行号 | set_index("日期") 后再 plot |
| 柱状图标签重叠 | 分类多 | rot=45 或改 barh |
| 保存图片模糊 | 分辨率低 | dpi=150 |
| 散点图无关联感 | 看不出关系 | 尝试 log 缩放或 hexbin |
14. 本章小结与练习
小结
- 入口:
df.plot(kind=...),kind决定图表类型; - 折线(line)、柱状(bar/barh)、直方(hist)、箱线(box)、散点(scatter)、饼图(pie)、面积(area)、密度(kde);
.plot()返回 Matplotlib Axes,可继续用 Matplotlib API 美化;ax=参数实现子图布局;- 中文显示需设置字体;时间序列先
set_index。
练习题
1. 对每日销量画折线图并叠加 7 日均线。
2. 用柱状图展示各区域销量,横向版本再画一遍。
3. 用散点图观察"广告费 vs 销量"关系并尝试解释。
4. 用箱线图比较两个区域的销量分布。
5. 制作一个 2×2 子图报告:趋势、分布、占比、关系。
下一篇预告:第 16 章 性能优化与实战案例 —— 让 pandas 跑得更快,并用一个完整项目串起全部知识。
文章回复
0 条公开回复