R / Richie全部文章 ↑

Python · 4 分钟阅读

Python:读取目录下的所有内容

目录


现代 Python 强烈推荐 pathlib.Path 来处理目录和文件。它跨平台、面向对象、并且
对 with open(...) 也很友好。本章以 pathlib 为主线,附带 os.walk 与
os.scandir 的对照写法。


1. 三种主流方式

方式 特点
pathlib.Path.iterdir / glob / rglob 跨平台、API 直观,推荐
os.scandir 速度最快,返回 DirEntry 含 stat 信息
os.walk 经典递归遍历,支持 topdown / followlinks

新代码优先用 pathlib。要“最高性能”或“需要 DirEntry 元信息”时用 os.scandir。


2. pathlib 推荐写法

2.1 列出目录里所有项

from pathlib import Path

p = Path("example_dir")

for item in p.iterdir():
    if item.is_file():
        print("file:", item)
    elif item.is_dir():
        print("dir :", item)

iterdir() 只列 直接子项,不递归。

2.2 递归列出所有文件

from pathlib import Path

for path in Path("example_dir").rglob("*"):   # ** 表示跨级
    print(path)

2.3 按扩展名过滤

from pathlib import Path

for txt in Path("example_dir").rglob("*.txt"):
    print(txt)

rglob 内部就是 Path.glob("**/pattern") + 递归。文件名以 . 开头的
文件默认会被 glob 排除(这是 POSIX shell 的传统行为)。

2.4 复制 / 移动 / 删除

from pathlib import Path

src = Path("a.txt")
dst = Path("b/a_copy.txt")
dst.parent.mkdir(parents=True, exist_ok=True)    # 自动建子目录
src.replace(dst)                                  # 重命名 / 移动

# 递归删除
import shutil
shutil.rmtree(Path("trash"), ignore_errors=True)

# 单文件删除
Path("a.txt").unlink(missing_ok=True)

3. 递归遍历(自己实现)

from pathlib import Path


def list_all(directory: Path):
    """打印目录下所有文件(递归)。"""
    for item in directory.iterdir():
        if item.is_file():
            print("file:", item)
        elif item.is_dir():
            print("dir :", item)
            list_all(item)            # 递归

Python 的递归深度默认 ~1000,超深目录会 RecursionError。目录级别深时
用 os.walk 或生成器版本:

from pathlib import Path
from typing import Iterator


def walk_all(root: Path) -> Iterator[Path]:
    """生成器:递归 yield 所有文件路径。"""
    for item in root.iterdir():
        if item.is_file():
            yield item
        elif item.is_dir():
            yield from walk_all(item)

4. os.walk(经典写法)

import os

for root, dirs, files in os.walk("example_dir"):
    print("当前目录:", root)
    print("子目录  :", dirs)
    print("文件    :", files)

要点:

  • 默认 topdown=True(先父后子)。topdown=False 时先访问子再父。

  • 修改 dirs 列表可以 剪枝:

    for root, dirs, files in os.walk("example_dir"):
        dirs[:] = [d for d in dirs if not d.startswith(".")]  # 跳过隐藏目录
        for f in files:
            print(root, f)
    
  • followlinks=False(默认)不会跟随软链,避免环。


5. os.scandir(性能最佳)

import os

with os.scandir("example_dir") as it:
    for entry in it:
        info = entry.stat()
        print(f"{entry.name:30s} size={info.st_size}  is_dir={entry.is_dir()}")
  • DirEntry.is_file() / is_dir() 不会 触发额外的 stat()。
  • 适合处理 超大型目录(百万级文件)。

6. 路径处理

from pathlib import Path

p = Path("/usr/local/bin/python3")

# 拆分
print(p.parts)     # ('/', 'usr', 'local', 'bin', 'python3')
print(p.name)      # python3
print(p.parent)    # /usr/local/bin
print(p.stem)      # python
print(p.suffix)    # .3
print(p.suffixes)  # ['.3']    注意:python3 视作无后缀的 "python3"

# 拼接(强烈优于字符串拼接)
new = p.parent / "share" / "doc"   # /usr/local/share/doc

跨平台分隔符 os.path.join 也能用,但 pathlib 更直观。


7. 实用脚本:批量处理文件

from pathlib import Path


def find_large_files(root: Path, min_size: int = 10 * 1024 * 1024):
    """找出 root 下大于 10 MiB 的文件。"""
    for path in root.rglob("*"):
        if path.is_file() and path.stat().st_size >= min_size:
            yield path


if __name__ == "__main__":
    for f in find_large_files(Path.home()):
        print(f"{f.stat().st_size:>12,d}  {f}")

8. 常见问题

  • iterdir() 顺序不稳定? 文件系统返回顺序不一定有序,需要顺序时显式
    sorted(p.iterdir())。
  • 路径不存在 / 没权限? iterdir 在路径不存在时抛 FileNotFoundError。
    先 if p.exists() 或 try / except。
  • 遍历时遍历到 . / ..? 不会,iterdir 自动排除。
  • 符号链接? 默认 不 跟随;想跟随用 Path.glob("**/*") 后再判断
    is_file(),或在 os.walk 中传 followlinks=True。
  • 中文 / 空格路径? pathlib 内部用 str,open 时显式传 encoding。

9. 最佳实践清单

  • 优先用 pathlib.Path,避免手写 os.path.join。
  • 任何文件操作都用 with 块。
  • 处理文本时显式传 encoding="utf-8"。
  • 遍历大目录优先 os.scandir / 生成器版本,避免一次性 list 全部。
  • 路径拼接用 / 运算符(Path / "subdir" / "file"),不要 f-string 拼字符串。
  • 删除 / 移动前先 print 确认,防止误操作。