本篇文章深入讲解 pandas 的一维数据结构 Series:如何创建、如何访问数据、有哪些常用属性和方法、如何参与运算。Series 是 DataFrame 的"积木",掌握它是学好 pandas 的第一步。
目录
1. 什么是 Series
2. 创建 Series 的 5 种方式
3. 索引(index)体系
4. 常用属性
5. 常用方法
6. Series 的运算
7. 缺失值处理入门
8. Series 的排序
9. name 属性与 Series 拼接
10. 常见坑与注意事项
11. 本章小结与练习
1. 什么是 Series
Series 是 pandas 的一维带标签数组,可以同时容纳:
- 数据值(values):任意类型(整数、浮点、字符串、对象等),通常为同质类型;
- 索引(index):每个值的标签,默认是
0, 1, 2, ...,也可以自定义。
可以用"带行号的一列数据"来理解它。
import pandas as pd
s = pd.Series([10, 20, 30])
print(s)
0 10
1 20
2 30
dtype: int64
左边的 0 1 2 是索引,右边的 10 20 30 是值,最后一行是数据类型。
2. 创建 Series 的 5 种方式
2.1 从列表 / 元组创建(最常用)
import pandas as pd
s1 = pd.Series([1, 2, 3, 4])
print(s1)
# 0 1
# 1 2
# 2 3
# 3 4
# dtype: int64
# 指定自定义索引
s2 = pd.Series([1, 2, 3], index=["a", "b", "c"])
print(s2)
# a 1
# b 2
# c 3
# dtype: int64
2.2 从字典创建(键自动成为索引)
s3 = pd.Series({"语文": 90, "数学": 95, "英语": 88})
print(s3)
# 语文 90
# 数学 95
# 英语 88
# dtype: int64
字典的键会按插入顺序成为索引(Python 3.7+ 保证字典有序)。
2.3 从 NumPy 数组创建
import numpy as np
arr = np.array([5.5, 6.6, 7.7])
s4 = pd.Series(arr)
print(s4)
# 0 5.5
# 1 6.6
# 2 7.7
# dtype: float64
# 注意:此时 Series 与 arr 共享内存,修改 arr 会影响 Series
arr[0] = 100
print(s4)
# 0 100.0
# ...
坑:从 ndarray 创建时默认共享底层数据;如果不希望共享,请用 s = pd.Series(arr.copy())。2.4 从标量创建(需要指定 index)
# 用同一个值填充多个位置
s5 = pd.Series(0, index=["a", "b", "c", "d"])
print(s5)
# a 0
# b 0
# c 0
# d 0
# dtype: int64
2.5 从 range / 生成器 / 其他 Series 创建
s6 = pd.Series(range(1, 6)) # 从 range
print(s6)
s7 = pd.Series(x * 2 for x in range(5)) # 从生成器
print(s7)
s8 = pd.Series(s3) # 从已有 Series(拷贝)
print(s8)
创建方式速查表
| 数据源 | 写法 | 索引来源 |
|--------|------|----------|
| 列表/元组 | pd.Series([1,2,3]) | 默认 RangeIndex |
| 字典 | pd.Series({"a":1}) | 字典键 |
| ndarray | pd.Series(np.array([...])) | 默认 RangeIndex |
| 标量 | pd.Series(0, index=["a","b"]) | 手动指定 |
| range/生成器 | pd.Series(range(5)) | 默认 RangeIndex |
| 已有 Series | pd.Series(s) | 继承原索引 |
3. 索引(index)体系
3.1 索引的类型
| 索引类型 | 说明 | 例子 |
|----------|------|------|
| RangeIndex | 默认的整数索引 | 0,1,2,... |
| Index | 普通标签索引 | 字符串等 |
| DatetimeIndex | 日期索引 | 2025-01-01, ... |
| MultiIndex | 多层索引 | (2025, "华东"), ... |
3.2 查看与修改索引
s = pd.Series([10, 20, 30], index=["x", "y", "z"])
# 查看索引
print(s.index) # Index(['x', 'y', 'z'], dtype='object')
# 重新赋值索引
s.index = ["a", "b", "c"]
print(s)
# 重置为默认索引(0,1,2...)
s2 = s.reset_index()
print(s2)
# index 0
# 0 a 10
# 1 b 20
# 2 c 30
3.3 索引的特性
- 索引可以重复(虽然通常不推荐);
- 索引可以是任意不可变类型:整数、字符串、日期、元组等;
- 索引用于对齐运算(见第 6 节),这是 pandas 最强大的特性之一。
# 索引可以重复
s_dup = pd.Series([1, 2, 3], index=["a", "a", "b"])
print(s_dup["a"]) # 返回一个 Series(多个匹配)
4. 常用属性
import pandas as pd
s = pd.Series([10, 20, 30, 40], index=["a", "b", "c", "d"], name="成绩")
# 属性一览
print(s.values) # 底层数组:[10 20 30 40](numpy 数组)
print(s.index) # 索引对象
print(s.dtype) # 数据类型:int64
print(s.shape) # 形状:(4,)
print(s.size) # 元素个数:4
print(s.ndim) # 维度:1
print(s.name) # 名称:成绩
print(s.nbytes) # 底层字节数
print(s.empty) # 是否为空:False
print(s.hasnans) # 是否含 NaN:False
print(s.is_unique) # 值是否唯一:True
| 属性 | 含义 |
|------|------|
| values | 底层 NumPy 数组 |
| index | 索引对象 |
| dtype | 元素数据类型 |
| shape / size / ndim | 形状 / 元素数 / 维度 |
| name | Series 的名称(列名) |
| empty | 是否为空 |
| hasnans | 是否含缺失值 |
| is_unique | 值是否全部唯一 |
5. 常用方法
s = pd.Series([3, 1, 2, 3, 4, 2, None])
# 描述性统计
print(s.describe()) # 数值型汇总统计
print(s.count()) # 非空个数
print(s.sum()) # 求和(自动跳过 NaN)
print(s.mean()) # 均值
print(s.median()) # 中位数
print(s.min(), s.max()) # 最小 / 最大值
print(s.std(), s.var()) # 标准差 / 方差
# 唯一值与计数
print(s.unique()) # 去重后的值数组 [3. 1. 2. 4. nan]
print(s.nunique()) # 唯一值个数(默认忽略 NaN):4
print(s.value_counts()) # 每个值的出现次数
value_counts() 输出示例:
2.0 2
3.0 2
1.0 1
4.0 1
Name: count, dtype: int64
其他常用方法:
s.head(2) # 前 2 个
s.tail(3) # 后 3 个
s.sample(2) # 随机抽 2 个
s.astype(str) # 类型转换
s.copy() # 拷贝
s.to_list() # 转 Python 列表
s.to_dict() # 转字典 {索引: 值}
6. Series 的运算
6.1 标量运算(向量化)
s = pd.Series([1, 2, 3])
print(s + 10) # 11 12 13
print(s * 2) # 2 4 6
print(s ** 2) # 1 4 9
print(s > 2) # False False True (返回布尔 Series)
6.2 Series 与 Series 运算:索引自动对齐
这是 pandas 最核心的特性:两个 Series 运算时,会按索引自动对齐,结果索引取并集,缺失位置填 NaN。
a = pd.Series([1, 2, 3], index=["x", "y", "z"])
b = pd.Series([10, 20, 30], index=["y", "z", "w"])
print(a + b)
# x NaN
# y 22.0
# z 33.0
# w NaN
# dtype: float64
x只在a中 → NaNw只在b中 → NaNy、z两边都有 → 正常相加
理解对齐(alignment) 是 pandas 进阶的第一道门槛。它让你无需手动"对上位置",但也要警惕意外产生 NaN。
6.3 算术方法与 fill_value
# 使用 add/sub/mul/div 等方法可以指定缺失值处理
print(a.add(b, fill_value=0))
# x 1.0
# y 22.0
# z 33.0
# w 30.0
# dtype: float64
常用算术方法:add、sub、mul、div、pow,均支持 fill_value 参数。
6.4 比较与逻辑运算
s = pd.Series([1, 5, 3, 8])
print(s > 3) # 布尔 Series
print((s > 1) & (s < 6)) # 与(注意用 & 而非 and)
print((s < 2) | (s > 7)) # 或(用 | 而非 or)
print(~(s > 3)) # 非(用 ~)
坑:pandas 中逻辑运算必须用&、|、~,不能使用 Python 的and、or、not(会报ValueError)。
7. 缺失值处理入门
import numpy as np
s = pd.Series([1, np.nan, 3, None, 5])
# 检测缺失
print(s.isna()) # False True False True False
print(s.isnull()) # 等价于 isna()
print(s.notna()) # 取反
# 删除缺失值
print(s.dropna())
# 0 1.0
# 2 3.0
# 4 5.0
# dtype: float64
# 填充缺失值
print(s.fillna(0))
# 0 1.0
# 1 0.0
# 2 3.0
# 3 0.0
# 4 5.0
# dtype: float64
# 前向填充 / 后向填充
print(s.ffill()) # 用上一个非空值填充
print(s.bfill()) # 用下一个非空值填充
缺失值专题将在第 7 章深入讲解。
8. Series 的排序
s = pd.Series([3, 1, 2], index=["c", "a", "b"])
# 按值排序(默认升序)
print(s.sort_values())
# a 1
# b 2
# c 3
# 按值降序
print(s.sort_values(ascending=False))
# 按索引排序
print(s.sort_index())
# 排名(从小到大排名次)
print(s.rank())
# c 3.0
# a 1.0
# b 2.0
9. name 属性与 Series 拼接
s1 = pd.Series([1, 2], name="a")
s2 = pd.Series([3, 4], name="b")
# 查看名字
print(s1.name) # a
# 拼接成 DataFrame(按列)
df = pd.concat([s1, s2], axis=1)
print(df)
# a b
# 0 1 3
# 1 2 4
10. 常见坑与注意事项
| 坑 | 现象 | 解决办法 |
|----|------|----------|
| 索引不匹配产生 NaN | 两个 Series 相加结果莫名多出 NaN | 理解对齐机制,用 fill_value 或先 reindex |
| 用 and/or 做逻辑运算 | ValueError: The truth value of a Series is ambiguous | 使用 &、\|、~,并加括号 |
| 从 ndarray 创建共享内存 | 修改原数组导致 Series 变化 | 用 .copy() |
| 忘记 name | DataFrame 的列名显示为 0 | 创建时指定 name= |
| 混合类型 | 整数+字符串 会变成 object 类型 | 尽量保持列类型单一 |
11. 本章小结与练习
小结
- Series 是 pandas 的一维带标签数组,由 values + index + name 组成;
- 5 种创建方式:列表、字典、ndarray、标量、range/生成器;
- 索引自动对齐是 pandas 运算的核心机制;
- 常用方法:
describe、value_counts、unique、sort_values、fillna、dropna等; - 逻辑运算用
&、|、~。
练习题
1. 用字典创建 Series:{"苹果": 5, "香蕉": 8, "橙子": 3},并打印。
2. 找出该 Series 中值大于 4 的元素。
3. 将 pd.Series([1,2,3]) 与 pd.Series([10,20,30], index=[1,2,3]) 相加,观察结果并解释为什么。
4. 对 Series [85, 92, None, 78, 95] 分别执行:求非空个数、均值、填充为 0 后的均值。
5. 用 value_counts 统计 ["男","女","男","男","女","女","女"] 的性别分布。
下一篇预告:第 3 章 DataFrame 详解 —— 掌握 pandas 最重要的二维表格数据结构。
文章回复
0 条公开回复