数据类型(dtype)决定了 pandas 怎么存储和计算数据:字符串列不能求和、object 列占内存大、日期列才能做时间运算。本篇文章系统讲解 pandas 的类型体系与各种转换武器。
目录
1. pandas 的数据类型体系
2. 查看类型:dtypes
3. 强制转换:astype
4. 智能转换:to_numeric / to_datetime / to_timedelta
5. category 分类类型
6. string 类型与 StringDtype
7. infer_objects 自动推断
8. 类型转换实战案例
9. 常见坑与注意事项
10. 本章小结与练习
1. pandas 的数据类型体系
1.1 常用 dtype 一览
| dtype | 说明 | 示例 |
|-------|------|------|
| int64 | 64 位整数 | 1, 2, 3 |
| float64 | 64 位浮点数 | 1.5, 2.0 |
| bool | 布尔 | True / False |
| object | Python 对象(字符串、混合类型) | "abc", "北京" |
| datetime64[ns] | 日期时间 | 2025-01-01 08:00:00 |
| timedelta64[ns] | 时间差 | 2 days 03:00:00 |
| category | 分类类型(有限取值) | 男/女, 低/中/高 |
| string | pandas 专用字符串类型 | "abc" |
| Int64/Float64 | 可含缺失值的整数/浮点 | 1, <NA> |
| complex | 复数 | 1+2j |
1.2 内存占用差异
import pandas as pd
import numpy as np
# int64 vs int32 内存差一倍
s64 = pd.Series(np.random.randint(0, 100, 1_000_000))
s32 = s64.astype("int32")
print(s64.memory_usage()) # 8,000,128 字节(约 8MB)
print(s32.memory_usage()) # 4,000,128 字节(约 4MB)
# object 字符串内存开销巨大
s_obj = pd.Series(["华东"] * 1_000_000)
print(s_obj.memory_usage()) # object 每元素存指针
print(s_obj.astype("category").memory_usage()) # category 大幅压缩
数据类型选择直接影响内存,海量数据场景差异可达数倍。
2. 查看类型:dtypes
df = pd.DataFrame({
"整数": [1, 2, 3],
"浮点": [1.5, 2.5, 3.5],
"字符串": ["a", "b", "c"],
"日期": pd.to_datetime(["2025-01-01", "2025-01-02", "2025-01-03"]),
"逻辑": [True, False, True],
})
# 每列类型
print(df.dtypes)
# 整数 int64
# 浮点 float64
# 字符串 object
# 日期 datetime64[ns]
# 逻辑 bool
# 筛选数值列 / 分类列 / 日期列
print(df.select_dtypes(include="number")) # 整数+浮点列
print(df.select_dtypes(include=["int64", "float64"]))
print(df.select_dtypes(include="object"))
print(df.select_dtypes(include="datetime64"))
print(df.select_dtypes(exclude="number")) # 排除数值列
2.1 单个 Series 的类型
s = pd.Series([1, 2, 3])
print(s.dtype) # dtype('int64')
print(s.dtype.name) # 'int64'
3. 强制转换:astype
astype() 是最直接的转换函数,能转则转,不能转就报错。
# 数值 -> 字符串
s = pd.Series([1, 2, 3])
print(s.astype(str))
# 0 1
# 1 2
# 2 3
# dtype: object
# 字符串 -> 数值
s2 = pd.Series(["1", "2", "3"])
print(s2.astype(int))
print(s2.astype(float))
# 转 bool
print(pd.Series([0, 1, 2]).astype(bool))
# 0 False
# 1 True
# 2 True
# DataFrame 批量转换
df = df.astype({"整数": float, "字符串": "string"})
print(df.dtypes)
# 全部列转换
df = df.astype(str)
astype 的局限:
# 报错:无法解析 "abc" 为整数
# pd.Series(["1", "abc"]).astype(int)
# ValueError: invalid literal for int() with base 10: 'abc'
# 报错:空字符串也无法转换
# pd.Series(["1", ""]).astype(int)
astype 是"严格模式":遇到无法解析的值直接抛异常。需要"宽容模式"用 to_numeric(errors="coerce")(见下节)。4. 智能转换:to_numeric / to_datetime / to_timedelta
4.1 to_numeric:字符串转数值
# 宽容模式:无法解析的变为 NaN
s = pd.Series(["1", "2.5", "abc", "3"])
print(pd.to_numeric(s, errors="coerce"))
# 0 1.0
# 1 2.5
# 2 NaN
# 3 3.0
# dtype: float64
# 直接对 DataFrame 某一列应用
df["金额"] = pd.to_numeric(df["金额"], errors="coerce")
# 批量:对所有 object 列尝试转换
for col in df.select_dtypes(include="object").columns:
df[col] = pd.to_numeric(df[col], errors="ignore")
| errors | 行为 |
|--------|------|
| "raise"(默认) | 遇到错误抛异常 |
| "coerce" | 无法解析的变为 NaN |
| "ignore" | 整列保持不变,返回原对象 |
4.2 to_datetime:转日期时间
# 各种日期字符串都能解析
print(pd.to_datetime("2025-01-01"))
print(pd.to_datetime("2025/01/01"))
print(pd.to_datetime("2025年1月1日"))
print(pd.to_datetime("2025-01-01 08:30:00"))
# 批量转换一列
df["日期"] = pd.to_datetime(df["日期"])
# 常见格式指定(解决无法自动识别的情况)
pd.to_datetime(df["日期"], format="%Y%m%d")
# 异常处理
pd.to_datetime(["2025-01-01", "not a date"], errors="coerce")
# DatetimeIndex(['2025-01-01', 'NaT'], dtype='datetime64[ns]', freq=None)
# 多列拼日期
df["日期"] = pd.to_datetime(df[["年", "月", "日"]])
# 带时区
pd.to_datetime("2025-01-01", utc=True)
4.3 to_timedelta:转时间差
# "2 days"、"3 hours" 等可解析
print(pd.to_timedelta("2 days 3 hours"))
print(pd.to_timedelta(["1 day", "2 days", "3 days"]))
# 与日期列运算
df["发货日期"] = df["下单日期"] + pd.to_timedelta(3, unit="D")
4.4 日期格式化输出
# datetime 转回字符串
df["日期字符串"] = df["日期"].dt.strftime("%Y-%m-%d")
5. category 分类类型
当一列只有少量固定取值(如性别、城市、级别)时,转成 category 有两个好处:省内存 + 支持有序比较。
# 转 category
s = pd.Series(["低", "中", "高", "低", "中"])
cat = s.astype("category")
print(cat.dtype) # CategoricalDtype(categories=['低', '中', '高'], ordered=False)
# 查看分类
print(cat.cat.categories) # Index(['低', '中', '高'], dtype='object')
# 有序分类(可以比较大小)
order = pd.CategoricalDtype(categories=["低", "中", "高"], ordered=True)
s_ordered = s.astype(order)
print(s_ordered > "低")
# 0 False
# 1 True
# 2 True
# 3 False
# 4 True
# 指定分类顺序后排序也会按定义顺序
print(s_ordered.sort_values())
# 0 低
# 3 低
# 1 中
# 4 中
# 2 高
典型应用:value_counts() 会按 categories 顺序显示,不会漏掉出现 0 次的分类。
city = pd.Series(["北京", "上海", "北京"]).astype(pd.CategoricalDtype(["北京", "上海", "广州", "深圳"]))
print(city.value_counts())
# 北京 2
# 上海 1
# 广州 0
# 深圳 0
6. string 类型与 StringDtype
pandas 1.0+ 提供了专门的 string 类型,替代 object 存储字符串:
s = pd.Series(["a", "b", None])
s_str = s.astype("string")
print(s_str.dtype) # string
# 好处
# 1. 缺失值用 <NA> 而不是 NaN
# 2. 字符串方法行为更一致
# 3. 内存占用可预测
# DataFrame 批量转 string
df = df.astype("string")
# 统计字符串列
print(s_str.str.upper())
# 0 A
# 1 B
# 2 <NA>
object vs string 对比:object 是"装任意 Python 对象的容器",string 是"专门的字符串类型"。新代码推荐用 astype("string")。7. infer_objects 自动推断
# 有时候读入后所有列都是 object,可以自动推断
df_raw = pd.DataFrame({
"a": ["1", "2", "3"], # 其实是数值
"b": ["2025-01-01", "2025-01-02", "2025-01-03"], # 其实是日期
})
df_inferred = df_raw.infer_objects()
print(df_inferred.dtypes)
# a object <- 注意:infer_objects 不推断字符串数字
# b object
# 更彻底的方式:convert_dtypes(pandas 1.0+)
df_conv = df_raw.convert_dtypes()
print(df_conv.dtypes)
# a string
# b string
convert_dtypes() 会把所有列转成 pandas 的"最佳"扩展类型(Int64、string、boolean 等),是快速规范化的一招。8. 类型转换实战案例
import pandas as pd
# 模拟从文件读入的"全 object"数据
raw = pd.DataFrame({
"订单号": ["A001", "A002", "A003"],
"金额": ["100.5", "200.0", "NaN"], # 文本数字 + NaN 标记
"数量": ["2", "3", "4"],
"下单日期": ["2025/1/1", "2025/1/2", "2025/1/3"],
"地区": ["华东", "华南", "华东"],
})
def normalize_types(df):
df = df.copy()
# 1. 数值列:宽容转换 + 缺失标记
df["金额"] = pd.to_numeric(df["金额"].replace("NaN", None), errors="coerce")
df["数量"] = pd.to_numeric(df["数量"], errors="coerce").astype("Int64")
# 2. 日期列
df["下单日期"] = pd.to_datetime(df["下单日期"])
# 3. 分类列
df["地区"] = df["地区"].astype("category")
# 4. 订单号保持字符串
df["订单号"] = df["订单号"].astype("string")
return df
df = normalize_types(raw)
print(df.dtypes)
print(df)
输出:
订单号 string
金额 float64
数量 Int64
下单日期 datetime64[ns]
地区 category
dtype: object
订单号 金额 数量 下单日期 地区
0 A001 100.5 2 2025-01-01 华东
1 A002 200.0 3 2025-01-02 华南
2 A003 NaN 4 2025-01-03 华东
9. 常见坑与注意事项
| 坑 | 现象 | 解决办法 |
|----|------|----------|
| astype 遇到坏数据报错 | ValueError | 用 to_numeric(errors="coerce") |
| int 列有 NaN | 转 int 报错 | 转 "Int64"(可空整数) |
| 中文数字/千分位 | 转数值失败 | 先清洗 str.replace(",", "") |
| 日期格式不统一 | 解析部分失败 | format= 指定 + errors="coerce" |
| object 列占内存 | 数据大内存爆 | 转 category / string |
| 转完后忘记赋值 | 原表没变 | df["列"] = df["列"].astype(...) |
10. 本章小结与练习
小结
- 常用类型:int64、float64、object、datetime64、category、string、Int64;
astype():严格模式强制转换;to_numeric(errors="coerce"):宽容转数值,坏值变 NaN;to_datetime(format=, errors=):日期解析;category:省内存 + 有序比较 + 完整 value_counts;convert_dtypes():一键最佳类型推断。
练习题
1. 将 ["1", "2.5", "abc", ""] 转成数值,要求坏值变 NaN。
2. 将 ["20250101", "20250102", "bad"] 转成日期,格式为 %Y%m%d。
3. 把性别列转成 category 并统计分布,验证 0 次出现的类别不丢失。
4. 对含 NaN 的整数列使用 "Int64" 转换并观察结果。
5. 用 convert_dtypes() 处理一份全 object 的数据并打印新类型。
下一篇预告:第 9 章 文本数据处理 —— 用 str 访问器玩转字符串:拆分、拼接、正则、清洗。
文章回复
0 条公开回复