就在前两周,我在工作室整理一批网页语料,脚本没有网络错误,JSON 文件也能正常打开,偏偏取字段时报了这么一行:
TypeError: list indices must be integers or slices, not str
桌上那杯美式已经凉了,我盯着代码看了几分钟,最后发现问题很普通:我以为 JSON 最外层是字典,实际拿到的是列表。
就这。
JSON 处理最容易让人误判的地方就在这里——“文件读取成功”和“数据提取正确”是两回事。下面不按语法手册讲,我直接用实际处理流程拆。
第一步不是取值,是看结构
假设有一个 records.json:
[
{
"id": 101,
"title": "Python 数据采集",
"author": {
"name": "沈砚",
"city": "杭州"
},
"tags": ["Python", "JSON"],
"status": "published"
},
{
"id": 102,
"title": "网页去重方法",
"author": null,
"tags": [],
"status": "draft"
}
]
先读取:
import json
with open("records.json", "r", encoding="utf-8") as file:
data = json.load(file)
print(type(data))
结果是:
<class 'list'>
所以这段代码是错的:
print(data["title"])
data 是列表,列表只能先用数字下标或循环访问:
print(data[0]["title"])
更常见的处理方式是遍历:
for item in data:
print(item.get("title"))
碰到 list indices must be integers,先别去搜索一堆复杂解释。打印 type(data),大概率马上就能看出问题。
字段不是每条都有
往下处理时,第二条记录的 author 是 null。
JSON 中的 null 会转换成 Python 的 None,所以这段代码会报错:
author_name = item.get("author", {
}).get("name")
乍看没问题,问题恰恰出在这儿:get("author", {}) 只会在键不存在时返回空字典,键存在且值为 None 时,返回的仍然是 None。
稳一点写:
author = item.get("author") or {
}
author_name = author.get("name", "未知作者")
完整遍历代码如下:
for item in data:
author = item.get("author") or {
}
title = item.get("title", "无标题")
author_name = author.get("name", "未知作者")
status = item.get("status", "unknown")
print(title, author_name, status)
输出:
Python 数据采集 沈砚 published
网页去重方法 未知作者 draft
我做语料这几年,对外部 JSON 有个不太乐观的默认判断:字段会缺、类型会变、原本应该是字典的位置也可能突然给你一个空字符串。听着有点悲观,只不过,这种悲观能让脚本少在凌晨报警。
按条件提取需要的数据
例如,只保留已经发布的记录:
published_records = [
item
for item in data
if item.get("status") == "published"
]
只提取标题:
titles = [
item.get("title")
for item in data
if item.get("title")
]
提取所有标签并去重:
all_tags = set()
for item in data:
tags = item.get("tags") or []
for tag in tags:
all_tags.add(tag)
print(list(all_tags))
这里没有直接写 item["tags"],原因还是一样:采集结果里出现缺字段很正常。程序应该决定缺字段时怎么处理,而不是把这个决定交给异常堆栈。
多层 JSON 不要硬接一串 get
层级一深,很多人会写成这样:
name = (
data.get("result", {
})
.get("user", {
})
.get("profile", {
})
.get("name")
)
两三层还能看,再深就开始费眼睛。我一般会封装一个提取函数,同时兼容字典和列表:
def deep_get(data, path, default=None):
current = data
for key in path:
if isinstance(current, dict):
if key not in current:
return default
current = current[key]
elif isinstance(current, list) and isinstance(key, int):
if key < 0 or key >= len(current):
return default
current = current[key]
else:
return default
return default if current is None else current
使用方式:
first_title = deep_get(
data,
[0, "title"],
"无标题"
)
author_city = deep_get(
data,
[0, "author", "city"],
"未知城市"
)
print(first_title)
print(author_city)
输出:
Python 数据采集
杭州
这个函数不复杂,它只是把“当前是字典就按键取,当前是列表就按下标取,结构不符合预期就返回默认值”集中处理了。项目里的 JSON 结构固定后,这种小函数比到处散落的异常判断好维护得多。
文件可能根本不是合法 JSON
读取文件时,至少把格式错误和文件错误接住:
import json
from pathlib import Path
def read_json(file_path):
path = Path(file_path)
try:
with path.open("r", encoding="utf-8-sig") as file:
return json.load(file)
except FileNotFoundError:
print(f"文件不存在:{path}")
except json.JSONDecodeError as error:
print(
f"JSON 格式错误:第 {error.lineno} 行,"
f"第 {error.colno} 列,{error.msg}"
)
except UnicodeDecodeError:
print("文件编码无法按 UTF-8 解析")
except OSError as error:
print(f"文件读取失败:{error}")
return None
调用时不要忘了判断:
data = read_json("records.json")
if data is None:
raise SystemExit("数据读取失败,程序终止")
JSONDecodeError 给出的行号和列号很有用。常见原因无非几种:用了单引号、末尾多了逗号、文件中带注释,或者把多个 JSON 对象一行一个地拼在了一起。
最后一种严格讲通常是 JSON Lines,也就是 .jsonl,不能用一次 json.load() 处理。
JSONL 要逐行读取
JSONL 文件通常长这样:
{"id": 101, "title": "Python 数据采集"}
{"id": 102, "title": "网页去重方法"}
{"id": 103, "title": "语料清洗"}
正确读取方式:
import json
records = []
with open("records.jsonl", "r", encoding="utf-8") as file:
for line_number, line in enumerate(file, start=1):
line = line.strip()
if not line:
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError as error:
print(
f"第 {line_number} 行解析失败:{error.msg}"
)
print(records)
注意这里用的是 json.loads(),因为循环里拿到的 line 是字符串。
json.load() 读文件,json.loads() 解字符串。名字只差一个 s,实际写脚本时混一次,半小时就没了(别问我怎么知道的)。
我现在处理 JSON 的顺序
拿到陌生 JSON,我不会马上写字段提取,而是先做四件事:
- 打印顶层类型,确认是字典还是列表。
- 查看一条样本,确认嵌套层级和字段类型。
- 区分必填字段与可选字段,确定默认值。
- 判断文件是普通 JSON、JSONL,还是需要流式读取的大文件。
文件只有几 MB 时,json.load() 足够。达到几百 MB,甚至大到内存放不下,就别再整份加载,可以考虑用 ijson 之类的流式解析工具,一条一条处理。
真正稳定的 JSON 代码,不是把取值写得多漂亮,而是结构跟预期不一样时,它依然知道该怎么办。