Ansible Roles:Linux 健康巡检与 LVM 磁盘清理
更新日期:2026-09-28
源代码目录:/Users/qiaoning5/PycharmProjects/PythonProject
这份笔记保存了本次编写的两个 role 的完整源码与运行方法。下方源码是生成笔记时的快照;以后若修改项目文件,应重新同步笔记。
1. 文件结构
PythonProject/
├── health-inspector.yml
├── lvm-disk-cleanup.yml
└── roles/
├── health-inspector/
│ ├── defaults/main.yml
│ ├── tasks/main.yml
│ └── files/collect.py
└── lvm-disk-cleanup/
├── defaults/main.yml
├── tasks/main.yml
└── files/cleanup.py
两个 Playbook 都使用 become: true。目标机需要 Linux 和 Python 3;控制机需要 Ansible。
2. Linux 健康巡检 role
覆盖指标
| 类别 | 指标 |
|---|---|
| 硬件 | 硬盘 SMART 5/187/197/198、温度、ECC CE/UCE、电源和风扇传感器 |
| 系统资源 | 1/5/15 分钟负载与每 CPU 负载、运行队列、I/O wait、MemAvailable、Swap、si/so、OOM、磁盘空间和 inode、iostat 的 util/await/IOPS/队列长度 |
| 网络 | ESTABLISHED、TIME_WAIT、CLOSE_WAIT、网卡错误与丢包、指定端口监听 |
| 服务 | 指定服务 active/enabled、systemctl failed units |
| 安全 | last/lastb、认证失败日志、UID 0 和空密码账户、口令期限、SSH root/密码认证、SUID/SGID 列表 |
| 日志 | dmesg error/fail/warn、本次启动 journal error、core dump 文件 |
每项在 JSON 中有 value 与 status。状态为 ok、warning、critical 或 unknown。SMART、IPMI、iostat 等依赖可选工具;缺少工具或硬件接口会标记为 unknown,不会误报正常。CE/UCE、网卡错误等是累计计数;iostat 使用 1 秒采样值。
运行
cd /Users/qiaoning5/PycharmProjects/PythonProject
ansible-playbook -i inventory health-inspector.yml \
-e '{"health_inspector_services":["sshd","nginx"],"health_inspector_ports":[22,80,443]}'
配置在 roles/health-inspector/defaults/main.yml。可用 -e 覆盖服务、端口、阈值、扫描范围和报告目录。报告默认保存在控制端 health_reports/<主机名>.json,权限为 0600。目标机安装 smartmontools、ipmitool、sysstat 后可获取更完整的数据。报告含登录记录和账户信息,应限制访问。
3. LVM 磁盘清理 role
工作流程
- 使用
lsblk、pvs、lvs识别物理磁盘、PV、VG、LV 的关系。 - 默认只显示计划与可执行命令,不删除磁盘数据。
- 执行模式先核对目标整盘、挂载点、Swap、VG 跨盘情况和确认字符串。
- 执行前再次发现并比较计划指纹;发现变化则拒绝继续。
- 依次运行
lvremove -ff -y、vgremove -ff -y、pvremove -ff -y、wipefs -af。 -
always任务删除目标机/tmp/ansible-lvm-disk-cleanup.py。如目标机断线,Ansible 无法完成清理,恢复连接后可重新运行默认发现命令或单独删除该文件。
这项操作不可逆。删除 docker VG 会删除其全部 Docker LV 和数据。wipefs 清除磁盘签名,并不保证安全擦除磁盘上所有数据块。role 接受 /dev/disk/by-id/ 整盘路径;无稳定路径的 VirtIO 整盘可明确使用 /dev/vdX。它拒绝分区路径、已挂载磁盘、Swap、跨越未选择磁盘的 VG,以及执行前发生变化的计划。
默认只读发现
你在有 ip.txt 的目录运行:
ansible-playbook -i ip.txt /Users/qiaoning5/PycharmProjects/PythonProject/lvm-disk-cleanup.yml
默认值是 lvm_cleanup_disks: []、lvm_cleanup_execute: false、lvm_cleanup_confirm: ''。默认运行会上传临时脚本进行发现,最后删除临时脚本;不会调用删除命令。输出中的 inventory 显示 VG/PV/LV/磁盘关系,suggested_targets 表示通过安全检查的目标,blocked_targets 说明不能生成命令的原因。
执行单台主机的 /dev/vdb
仅在确认该主机的 /dev/vdb 是要清空的整盘、业务已停止并有备份后运行。以下命令只针对 11.50.138.135;其他主机替换 --limit 值,逐台执行与核对结果。
ansible-playbook -i ip.txt /Users/qiaoning5/PycharmProjects/PythonProject/lvm-disk-cleanup.yml \
--limit 11.50.138.135 \
-e '{"lvm_cleanup_disks":["/dev/vdb"],"lvm_cleanup_execute":true,"lvm_cleanup_confirm":"ERASE_LVM_AND_WIPEFS"}'
本次用户贴出的只读发现结果显示:11.50.138.134、11.50.138.135、11.50.139.1、11.50.139.122 的 /dev/vdb 属于 docker VG;11.50.139.26 没有发现 PV。该结果是历史快照,执行前仍以目标机当前发现结果为准。若使用了另一份 disk/lvm-disk-cleanup.yml,需确认其 role 源码已同步本笔记中的最新版本。
单独删除旧运行遗留的临时文件
ansible -i ip.txt all -b -m file -a 'path=/tmp/ansible-lvm-disk-cleanup.py state=absent'
这条命令只删除临时 Python 文件,不删除 LVM 或磁盘数据。
4. 完整源码快照
以下代码块按实际目录组织,复制到同名文件即可还原本次两个 role。
health-inspector.yml
---
- name: Linux 健康巡检
hosts: all
become: true
gather_facts: false
roles:
- health-inspector
roles/health-inspector/defaults/main.yml
---
health_inspector_report_dir: "{{ playbook_dir }}/health_reports"
health_inspector_services: []
health_inspector_ports: []
health_inspector_suid_paths: ['/usr/bin', '/usr/sbin', '/bin', '/sbin']
health_inspector_core_paths: ['/var/lib/systemd/coredump', '/var/crash', '/var/core']
health_inspector_log_lines: 30
health_inspector_thresholds:
smart_temperature_warn_c: 50
smart_temperature_critical_c: 60
load_per_cpu_warn: 0.7
load_per_cpu_critical: 1.0
iowait_warn_percent: 20
memory_available_warn_percent: 10
swap_used_warn_percent: 50
filesystem_warn_percent: 80
filesystem_critical_percent: 95
inode_warn_percent: 80
disk_util_warn_percent: 80
disk_await_warn_ms: 50
time_wait_warn: 5000
close_wait_warn: 100
roles/health-inspector/tasks/main.yml
---
- name: 创建控制端报告目录
ansible.builtin.file:
path: "{{ health_inspector_report_dir }}"
state: directory
mode: '0700'
delegate_to: localhost
become: false
- name: 上传只读巡检程序
ansible.builtin.copy:
src: collect.py
dest: /tmp/ansible-health-inspector-collect.py
mode: '0700'
- name: 采集健康指标
ansible.builtin.command:
argv:
- python3
- /tmp/ansible-health-inspector-collect.py
- "{{ health_inspector_services | to_json }}"
- "{{ health_inspector_ports | to_json }}"
- "{{ health_inspector_suid_paths | to_json }}"
- "{{ health_inspector_core_paths | to_json }}"
- "{{ health_inspector_thresholds | to_json }}"
- "{{ health_inspector_log_lines | int | string }}"
register: health_inspector_raw
changed_when: false
failed_when: health_inspector_raw.rc != 0
- name: 解析巡检结果
ansible.builtin.set_fact:
health_inspector_result: "{{ health_inspector_raw.stdout | from_json }}"
- name: 保存本地 JSON 报告
ansible.builtin.copy:
content: "{{ health_inspector_result | to_nice_json }}\n"
dest: "{{ health_inspector_report_dir }}/{{ inventory_hostname | regex_replace('[^A-Za-z0-9_.-]', '_') }}.json"
mode: '0600'
delegate_to: localhost
become: false
- name: 显示巡检汇总
ansible.builtin.debug:
msg:
host: "{{ inventory_hostname }}"
status: "{{ health_inspector_result.status }}"
critical: "{{ health_inspector_result.summary.critical }}"
warning: "{{ health_inspector_result.summary.warning }}"
unknown: "{{ health_inspector_result.summary.unknown }}"
report: "{{ health_inspector_report_dir }}/{{ inventory_hostname | regex_replace('[^A-Za-z0-9_.-]', '_') }}.json"
- name: 删除目标机临时巡检程序
ansible.builtin.file:
path: /tmp/ansible-health-inspector-collect.py
state: absent
changed_when: false
roles/health-inspector/files/collect.py
#!/usr/bin/env python3
"""Linux read-only health snapshot. Optional tools yield unknown, never healthy."""
import datetime
import json
import os
import re
import shutil
import subprocess
import sys
from pathlib import Path
def run(args, timeout=12):
if not shutil.which(args[0]):
return None
try:
p = subprocess.run(args, text=True, stdout=subprocess.PIPE,
stderr=subprocess.PIPE, timeout=timeout, check=False)
return p.stdout if p.returncode == 0 else None
except (OSError, subprocess.TimeoutExpired):
return None
def read(path):
try:
return Path(path).read_text(errors='replace')
except OSError:
return None
def number(value):
try:
return float(value)
except (TypeError, ValueError):
return None
def sample(text, limit):
return (text or '').splitlines()[-limit:]
def add(group, name, value, status='ok', detail=None):
item = {'value': value, 'status': status}
if detail is not None:
item['detail'] = detail
result['metrics'].setdefault(group, {})[name] = item
def graded(value, warn=None, critical=None, reverse=False):
if value is None:
return 'unknown'
if reverse:
return 'critical' if critical is not None and value <= critical else ('warning' if warn is not None and value <= warn else 'ok')
return 'critical' if critical is not None and value >= critical else ('warning' if warn is not None and value >= warn else 'ok')
def hardware():
devices = []
raw = run(['lsblk', '-J', '-d', '-o', 'NAME,TYPE,TRAN'])
if raw:
try:
devices = ['/dev/' + d['name'] for d in json.loads(raw)['blockdevices'] if d.get('type') == 'disk']
except (KeyError, ValueError, TypeError):
pass
for dev in devices:
# smartctl uses nonzero bitmask exit codes even when valid SMART data exists.
try:
raw = subprocess.run(['smartctl', '-j', '-a', dev], text=True,
capture_output=True, timeout=25).stdout if shutil.which('smartctl') else None
except (OSError, subprocess.TimeoutExpired):
raw = None
if raw is None:
add('hardware', dev + '.smart', None, 'unknown', 'smartctl 缺失、设备不支持或访问失败')
continue
try:
data = json.loads(raw)
except ValueError:
add('hardware', dev + '.smart', None, 'unknown', 'SMART JSON 无效')
continue
table = {x.get('id'): x for x in data.get('ata_smart_attributes', {}).get('table', [])}
for ident, label in [(5, 'Reallocated_Sector_Ct'), (187, 'Reported_Uncorrect'),
(197, 'Current_Pending_Sector'), (198, 'Offline_Uncorrectable')]:
v = number(table.get(ident, {}).get('raw', {}).get('value'))
add('hardware', dev + '.smart_' + str(ident) + '_' + label, v,
graded(v, 1, 1), '原始计数;NVMe 或不支持的设备为 unknown')
temp = number(data.get('temperature', {}).get('current'))
add('hardware', dev + '.temperature_c', temp,
graded(temp, th['smart_temperature_warn_c'], th['smart_temperature_critical_c']))
if not devices:
add('hardware', 'smart_devices', None, 'unknown', 'lsblk 未发现磁盘或不可用')
edac = Path('/sys/devices/system/edac/mc')
for suffix, label in [('ce_count', 'ecc_ce'), ('ue_count', 'ecc_uce')]:
vals = [number(read(p).strip()) for p in edac.glob('mc[0-9]*/' + suffix)] if edac.exists() else []
total = int(sum(vals)) if vals and all(v is not None for v in vals) else None
add('hardware', label, total, graded(total, 1, 1), '开机以来累计次数')
raw = run(['ipmitool', 'sdr', 'elist', 'all'], 20)
if raw is None:
add('hardware', 'power_supplies', None, 'unknown', '需要 IPMI/BMC 和 ipmitool')
add('hardware', 'fans', None, 'unknown', '需要 IPMI/BMC 和 ipmitool')
else:
for key, regex in [('power_supplies', r'(?i)\b(?:psu|power supply|pwr)\b'),
('fans', r'(?i)\bfan\b')]:
lines = [x for x in raw.splitlines() if re.search(regex, x)]
bad = [x for x in lines if re.search(r'(?i)\b(?:fail|critical|not present|no reading|nr|cr|nc)\b', x)]
incomplete_psu = key == 'power_supplies' and len(lines) < 2
add('hardware', key, lines or None, 'unknown' if not lines or incomplete_psu else ('critical' if bad else 'ok'),
'传感器原始状态;双路电源需核对两条 PSU 记录')
def resources():
load = os.getloadavg()
cpu = os.cpu_count() or 1
for index, interval in enumerate((1, 5, 15)):
add('resources', 'load_' + str(interval) + 'm', round(load[index], 2))
ratio = round(load[index] / cpu, 3)
add('resources', 'load_' + str(interval) + 'm_per_cpu', ratio,
graded(ratio, th['load_per_cpu_warn'], th['load_per_cpu_critical']))
raw = run(['vmstat', '1', '2'])
if raw:
rows = raw.splitlines()
cols = rows[-1].split() if len(rows) >= 4 else []
# Linux vmstat: r b swpd free buff cache si so bi bo in cs us sy id wa st
for key, pos, warn in [('run_queue', 0, cpu), ('swap_in_kb_s', 6, 1),
('swap_out_kb_s', 7, 1), ('cpu_iowait_percent', 15, th['iowait_warn_percent'])]:
v = number(cols[pos]) if len(cols) > pos else None
add('resources', key, v, graded(v, warn))
else:
for key in ('run_queue', 'swap_in_kb_s', 'swap_out_kb_s', 'cpu_iowait_percent'):
add('resources', key, None, 'unknown', 'vmstat 不可用')
mem = {m.group(1): int(m.group(2)) for m in re.finditer(r'^(\w+):\s+(\d+) kB', read('/proc/meminfo') or '', re.M)}
available = mem.get('MemAvailable')
pct = round(available / mem['MemTotal'] * 100, 2) if available is not None and mem.get('MemTotal') else None
add('resources', 'memory_available_kb', available, 'unknown' if available is None else 'ok')
add('resources', 'memory_available_percent', pct, graded(pct, th['memory_available_warn_percent'], reverse=True))
swap = mem.get('SwapTotal', 0)
used = swap - mem.get('SwapFree', 0)
add('resources', 'swap_used_kb', used)
add('resources', 'swap_used_percent', round(used / swap * 100, 2) if swap else 0,
graded(used / swap * 100 if swap else 0, th['swap_used_warn_percent']))
journal = run(['journalctl', '-k', '-b', '--no-pager', '-q'], 20)
oom = [x for x in (journal or '').splitlines() if re.search(r'(?i)oom-killer|out of memory|killed process', x)]
add('resources', 'oom_events_current_boot', sample('\n'.join(oom), limit) if journal is not None else None,
'unknown' if journal is None else graded(len(oom), 1, 1))
for mode, key, threshold in [([], 'filesystem', 'filesystem_warn_percent'),
(['-i'], 'inodes', 'inode_warn_percent')]:
raw = run(['df', '-P'] + mode)
if raw is None:
add('resources', key, None, 'unknown', 'df 不可用')
continue
items = []
for line in raw.splitlines()[1:]:
parts = line.split()
if len(parts) >= 6 and parts[4].endswith('%') and not parts[0].startswith(('tmpfs', 'devtmpfs')):
items.append({'mount': parts[-1], 'used_percent': int(parts[4][:-1])})
status = 'critical' if key == 'filesystem' and any(x['used_percent'] >= th['filesystem_critical_percent'] for x in items) else graded(max((x['used_percent'] for x in items), default=None), th[threshold])
add('resources', key, items, status)
raw = run(['iostat', '-dx', '-y', '1', '2'], 15)
if raw is None:
add('resources', 'disk_io', None, 'unknown', '需要 sysstat/iostat')
else:
lines = raw.splitlines()
headers = [i for i, x in enumerate(lines) if x.strip().startswith('Device')]
items = []
if headers:
header = lines[headers[-1]].split()
for line in lines[headers[-1] + 1:]:
cols = line.split()
if len(cols) != len(header):
continue
row = dict(zip(header, cols))
def field(*names):
return next((number(row[n]) for n in names if n in row), None)
rs, ws = field('r/s'), field('w/s')
item = {'device': cols[0], 'util_percent': field('%util'),
'await_ms': field('await'), 'iops': rs + ws if rs is not None and ws is not None else None,
'avgqu_sz': field('avgqu-sz', 'aqu-sz')}
items.append(item)
warn = any((x['util_percent'] or 0) >= th['disk_util_warn_percent'] or
(x['await_ms'] or 0) >= th['disk_await_warn_ms'] for x in items)
add('resources', 'disk_io', items or None, 'unknown' if not items else ('warning' if warn else 'ok'),
'采样区间 1 秒;iostat -y 排除开机以来平均值')
def network():
raw = run(['ss', '-tan'])
if raw is None:
add('network', 'tcp_states', None, 'unknown', 'ss 不可用')
else:
counts = {state: len(re.findall(r'^' + state + r'\s', raw, re.M)) for state in ('ESTAB', 'TIME-WAIT', 'CLOSE-WAIT')}
add('network', 'established', counts['ESTAB'])
add('network', 'time_wait', counts['TIME-WAIT'], graded(counts['TIME-WAIT'], th['time_wait_warn']))
add('network', 'close_wait', counts['CLOSE-WAIT'], graded(counts['CLOSE-WAIT'], th['close_wait_warn']))
nic = {}
for path in Path('/sys/class/net').glob('*'):
if path.name == 'lo':
continue
nic[path.name] = {}
for field in ('rx_errors', 'tx_errors', 'rx_dropped', 'tx_dropped'):
raw = read(path / 'statistics' / field)
nic[path.name][field] = int(raw.strip()) if raw and raw.strip().isdigit() else None
add('network', 'interfaces', nic or None,
'unknown' if not nic else ('warning' if any((v or 0) > 0 for fields in nic.values() for v in fields.values()) else 'ok'),
'开机以来累计计数')
listening = run(['ss', '-ltnH'])
for port in ports:
found = listening is not None and any(re.search(r':' + str(port) + r'\s', line) for line in listening.splitlines())
add('network', 'listen_' + str(port), found if listening is not None else None,
'unknown' if listening is None else ('ok' if found else 'critical'))
def services():
for svc in svc_names:
# systemctl returns nonzero for inactive/disabled; query output remains useful.
for key, command, expected in [('active', 'is-active', 'active'), ('enabled', 'is-enabled', 'enabled')]:
try:
p = subprocess.run(['systemctl', command, svc], capture_output=True, text=True, timeout=8)
val = p.stdout.strip() or 'unknown'
except (OSError, subprocess.TimeoutExpired):
val = 'unknown'
add('services', svc + '.' + key, val, 'unknown' if val == 'unknown' else ('ok' if val == expected else 'critical'))
raw = run(['systemctl', '--failed', '--no-legend', '--plain'])
failed = [x.strip() for x in (raw or '').splitlines() if x.strip() and not x.startswith('0 loaded')]
add('services', 'failed_units', failed if raw is not None else None,
'unknown' if raw is None else graded(len(failed), 1, 1))
def security():
for key, command in [('successful_logins', ['last', '-n', str(limit)]),
('failed_logins', ['lastb', '-n', str(limit)])]:
raw = run(command)
add('security', key, sample(raw, limit) if raw is not None else None,
'unknown' if raw is None else 'ok')
auth = next((read(p) for p in ('/var/log/secure', '/var/log/auth.log') if read(p) is not None), None)
failures = [x for x in (auth or '').splitlines() if re.search(r'Failed password|authentication failure', x, re.I)]
add('security', 'auth_failures', sample('\n'.join(failures), limit) if auth is not None else None,
'unknown' if auth is None else graded(len(failures), 1))
passwd = read('/etc/passwd')
shadow = read('/etc/shadow')
uid0 = [x.split(':')[0] for x in (passwd or '').splitlines() if len(x.split(':')) > 2 and x.split(':')[2] == '0']
empty = [x.split(':')[0] for x in (shadow or '').splitlines() if len(x.split(':')) > 1 and x.split(':')[1] == '']
add('security', 'uid0_accounts', uid0 if passwd is not None else None,
'unknown' if passwd is None else ('warning' if len(uid0) > 1 else 'ok'))
add('security', 'empty_password_accounts', empty if shadow is not None else None,
'unknown' if shadow is None else graded(len(empty), 1, 1))
defs = read('/etc/login.defs')
policy = {}
for key in ('PASS_MAX_DAYS', 'PASS_MIN_DAYS', 'PASS_WARN_AGE'):
m = re.search(r'^\s*' + key + r'\s+(\S+)', defs or '', re.M)
policy[key] = m.group(1) if m else None
add('security', 'password_age_policy', policy, 'unknown' if defs is None else 'ok',
'login.defs 默认值;已有账户策略可用 chage -l 单独核查')
ssh = run(['sshd', '-T']) or run(['/usr/sbin/sshd', '-T'])
for key, expected in [('permitrootlogin', 'no'), ('passwordauthentication', 'no')]:
m = re.search(r'^' + key + r'\s+(\S+)', ssh or '', re.M | re.I)
value = m.group(1) if m else None
add('security', 'ssh_' + key, value, 'unknown' if value is None else ('ok' if value == expected else 'warning'))
suid = []
for base in suid_paths:
if not os.path.isdir(base):
continue
for root, dirs, files in os.walk(base):
dirs[:] = [d for d in dirs if not os.path.islink(os.path.join(root, d))]
for name in files:
path = os.path.join(root, name)
try:
mode = os.stat(path, follow_symlinks=False).st_mode
if mode & 0o6000:
suid.append({'path': path, 'mode': oct(mode & 0o7777)})
except OSError:
pass
add('security', 'suid_sgid_files', suid, 'ok', '用于基线对比;存在 SUID/SGID 本身不判异常')
def logs():
raw = run(['dmesg', '--color=never'])
hits = [x for x in (raw or '').splitlines() if re.search(r'\b(error|fail|warn)\b', x, re.I)]
add('logs', 'dmesg_error_fail_warn', sample('\n'.join(hits), limit) if raw is not None else None,
'unknown' if raw is None else graded(len(hits), 1))
raw = run(['journalctl', '-b', '-p', 'err', '--no-pager', '-q', '-n', str(limit)], 20)
lines = [x for x in (raw or '').splitlines() if x.strip()]
add('logs', 'journal_current_boot_errors', lines if raw is not None else None,
'unknown' if raw is None else graded(len(lines), 1))
cores = []
for base in core_paths:
if os.path.isdir(base):
for root, dirs, files in os.walk(base):
for name in files:
path = os.path.join(root, name)
try:
st = os.stat(path)
cores.append({'path': path, 'size_bytes': st.st_size})
except OSError:
pass
add('logs', 'core_dump_files', cores, graded(len(cores), 1))
if __name__ == '__main__':
svc_names, ports, suid_paths, core_paths, th = [json.loads(x) for x in sys.argv[1:6]]
limit = max(1, min(200, int(sys.argv[6])))
result = {'host': os.uname().nodename, 'collected_at': datetime.datetime.now(datetime.timezone.utc).isoformat(), 'metrics': {}}
for collector in (hardware, resources, network, services, security, logs):
try:
collector()
except Exception as exc:
add('collection_errors', collector.__name__, str(exc), 'unknown')
summary = {s: sum(x['status'] == s for group in result['metrics'].values() for x in group.values())
for s in ('ok', 'warning', 'critical', 'unknown')}
result['summary'] = summary
result['status'] = 'critical' if summary['critical'] else ('warning' if summary['warning'] else ('unknown' if summary['unknown'] else 'ok'))
print(json.dumps(result, ensure_ascii=False))
lvm-disk-cleanup.yml
---
- name: 清理明确指定的 LVM 物理磁盘
hosts: all
become: true
gather_facts: false
roles:
- lvm-disk-cleanup
roles/lvm-disk-cleanup/defaults/main.yml
---
# 优先使用 /dev/disk/by-id;无稳定路径的 VirtIO 整盘可显式指定 /dev/vdX。
lvm_cleanup_disks: []
lvm_cleanup_confirm: ''
lvm_cleanup_execute: false
# 默认沿用本次运行的 inventory;也可显式覆盖。
lvm_cleanup_inventory_path: "{{ ansible_inventory_sources[0] }}"
roles/lvm-disk-cleanup/tasks/main.yml
---
- name: 运行 LVM 巡检与清理并复原临时环境
block:
- name: 上传 LVM 发现与清理程序
ansible.builtin.copy:
src: cleanup.py
dest: /tmp/ansible-lvm-disk-cleanup.py
mode: '0700'
- name: 发现 LVM 磁盘并执行安全检查
ansible.builtin.command:
argv:
- python3
- /tmp/ansible-lvm-disk-cleanup.py
- plan
- "{{ lvm_cleanup_disks | to_json }}"
register: lvm_cleanup_plan_raw
changed_when: false
- name: 导出清理计划
ansible.builtin.set_fact:
lvm_cleanup_plan: "{{ lvm_cleanup_plan_raw.stdout | from_json }}"
- name: 显示清理计划
ansible.builtin.debug:
var: lvm_cleanup_plan
- name: 输出已通过安全检查的完整执行命令
ansible.builtin.debug:
msg: >-
ansible-playbook -i {{ lvm_cleanup_inventory_path | quote }} {{ (playbook_dir ~ '/lvm-disk-cleanup.yml') | quote }}
--limit {{ inventory_hostname | quote }}
-e {{ {'lvm_cleanup_disks': [item.by_id], 'lvm_cleanup_execute': true,
'lvm_cleanup_confirm': 'ERASE_LVM_AND_WIPEFS'} | to_json | quote }}
loop: "{{ lvm_cleanup_plan.suggested_targets | default([]) }}"
loop_control:
label: "{{ item.disk }}"
when: not (lvm_cleanup_execute | bool)
- name: 提示无可直接清理的磁盘
ansible.builtin.debug:
msg: '未发现可生成执行命令的磁盘;可能有挂载、Swap、VG 跨盘或缺少 /dev/disk/by-id 路径。'
when:
- not (lvm_cleanup_execute | bool)
- lvm_cleanup_plan.suggested_targets | default([]) | length == 0
- name: 显示不能生成执行命令的具体原因
ansible.builtin.debug:
msg: "{{ item.disk }}: {{ item.reason }}"
loop: "{{ lvm_cleanup_plan.blocked_targets | default([]) }}"
loop_control:
label: "{{ item.disk }}"
when: not (lvm_cleanup_execute | bool)
- name: 仅在明确授权时执行清理
when: lvm_cleanup_execute | bool
block:
- name: 核对清理授权
ansible.builtin.assert:
that:
- lvm_cleanup_confirm == 'ERASE_LVM_AND_WIPEFS'
- lvm_cleanup_plan.disks | length > 0
fail_msg: '必须明确设置目标盘、lvm_cleanup_execute=true 和确认字符串。'
- name: 删除 LV/VG/PV 并擦除磁盘签名
ansible.builtin.command:
argv:
- python3
- /tmp/ansible-lvm-disk-cleanup.py
- execute
- "{{ lvm_cleanup_disks | to_json }}"
- "{{ lvm_cleanup_plan.fingerprint }}"
- "{{ lvm_cleanup_confirm }}"
register: lvm_cleanup_result
changed_when: true
- name: 显示执行结果
ansible.builtin.debug:
msg: "{{ lvm_cleanup_result.stdout | from_json }}"
always:
- name: 删除目标机临时巡检程序
ansible.builtin.file:
path: /tmp/ansible-lvm-disk-cleanup.py
state: absent
changed_when: false
roles/lvm-disk-cleanup/files/cleanup.py
#!/usr/bin/env python3
"""Discover LVM disk ownership; destroy only explicitly selected, idle disks."""
import hashlib
import json
import os
import re
import shutil
import subprocess
import sys
def command(argv):
proc = subprocess.run(argv, text=True, stdout=subprocess.PIPE,
stderr=subprocess.PIPE, timeout=60, check=False)
if proc.returncode:
raise RuntimeError('%s: %s' % (' '.join(argv), proc.stderr.strip()))
return proc.stdout
def report(command_name, columns):
data = json.loads(command([command_name, '--reportformat', 'json',
'--units', 'b', '--nosuffix', '-o', columns]))
return data['report'][0][{'pvs': 'pv', 'lvs': 'lv'}[command_name]]
def discover(requested, include_suggestions=True):
if not isinstance(requested, list) or any(not isinstance(x, str) for x in requested):
raise ValueError('lvm_cleanup_disks 必须是设备路径列表')
if len(requested) != len(set(requested)):
raise ValueError('目标磁盘重复')
for tool in ('lsblk', 'pvs', 'lvs', 'lvremove', 'vgremove', 'pvremove', 'wipefs'):
if shutil.which(tool) is None:
raise RuntimeError('缺少命令: ' + tool)
# MOUNTPOINTS (plural), SERIAL and WWN are absent on older util-linux.
tree = json.loads(command(['lsblk', '-J', '-p', '-o', 'NAME,TYPE,MOUNTPOINT']))['blockdevices']
by_path = {}
disks = {}
def walk(node, disk):
path = os.path.realpath(node['name'])
if node['type'] == 'disk':
disk = path
disks[disk] = {'path': path, 'nodes': []}
if disk is not None:
mountpoint = node.get('mountpoint')
mountpoints = [mountpoint] if mountpoint else []
disks[disk]['nodes'].append({'path': path, 'type': node['type'],
'mountpoints': mountpoints})
by_path[path] = disk
for child in node.get('children') or []:
walk(child, disk)
for root in tree:
walk(root, None)
pvs = report('pvs', 'pv_name,vg_name')
lvs = report('lvs', 'lv_path,vg_name')
pv_by_vg = {}
for pv in pvs:
vg = (pv.get('vg_name') or '').strip()
if vg:
pv_by_vg.setdefault(vg, []).append(os.path.realpath(pv['pv_name'].strip()))
lv_by_vg = {}
for lv in lvs:
vg = (lv.get('vg_name') or '').strip()
path = (lv.get('lv_path') or '').strip()
if vg and path:
lv_by_vg.setdefault(vg, []).append(path)
chosen = []
for name in requested:
stable = name.startswith('/dev/disk/by-id/')
virtio = re.fullmatch(r'/dev/vd[a-z]+', name) is not None
if not stable and not virtio:
raise ValueError('目标必须使用 /dev/disk/by-id/ 或明确的 /dev/vdX 整盘路径: ' + name)
if not os.path.exists(name):
raise ValueError('设备不存在: ' + name)
path = os.path.realpath(name)
if path not in disks:
raise ValueError('目标不是整块物理磁盘: ' + name)
if virtio and path not in {pv for paths in pv_by_vg.values() for pv in paths}:
raise ValueError('VirtIO 整盘自身不是已加入 VG 的 PV: ' + name)
if path in chosen:
raise ValueError('不同路径解析到同一磁盘: ' + path)
chosen.append(path)
swap_paths = set()
with open('/proc/swaps', encoding='utf-8') as stream:
for line in stream.readlines()[1:]:
swap_paths.add(os.path.realpath(line.split()[0]))
all_pv_disks = {}
for vg, paths in pv_by_vg.items():
all_pv_disks[vg] = [by_path.get(p) for p in paths]
targets = []
for disk in chosen:
nodes = disks[disk]['nodes']
if any(n['mountpoints'] for n in nodes):
raise ValueError('目标盘有挂载点: ' + disk)
if any(n['path'] in swap_paths for n in nodes):
raise ValueError('目标盘正在用于 Swap: ' + disk)
related = sorted(vg for vg, owners in all_pv_disks.items() if disk in owners)
if not related:
raise ValueError('目标盘没有已加入 VG 的 PV: ' + disk)
targets.append({'requested': requested[chosen.index(disk)], 'disk': disk,
'vgs': related})
target_vgs = sorted({vg for target in targets for vg in target['vgs']})
for vg in target_vgs:
owners = all_pv_disks[vg]
if any(owner is None or owner not in chosen for owner in owners):
raise ValueError('VG %s 跨越未选择或无法识别的磁盘;必须整体选择' % vg)
if any(os.path.realpath(lv) in swap_paths for lv in lv_by_vg.get(vg, [])):
raise ValueError('VG %s 中有正在使用的 Swap LV' % vg)
inventory = [{'vg': vg, 'pvs': pv_by_vg[vg], 'disks': all_pv_disks[vg],
'lvs': sorted(lv_by_vg.get(vg, []))} for vg in sorted(pv_by_vg)]
plan = {'disks': targets, 'vgs': target_vgs, 'inventory': inventory,
'lvs': {vg: sorted(lv_by_vg.get(vg, [])) for vg in target_vgs},
'pvs': {vg: sorted(pv_by_vg[vg]) for vg in target_vgs},
'discovered_pv_disks': sorted(set(x for owners in all_pv_disks.values() for x in owners if x))}
if include_suggestions and not requested:
suggestions = []
blocked = []
# Only stable whole-disk aliases, never partition aliases.
alias_dir = '/dev/disk/by-id'
aliases = {}
if os.path.isdir(alias_dir):
for name in sorted(os.listdir(alias_dir)):
if '-part' in name or not name.startswith(('wwn-', 'nvme-eui.', 'ata-', 'scsi-')):
continue
alias = os.path.join(alias_dir, name)
disk = os.path.realpath(alias)
if disk in disks and disk not in aliases:
aliases[disk] = alias
for disk in plan['discovered_pv_disks']:
alias = aliases.get(disk)
if alias:
try:
candidate = discover([alias], include_suggestions=False)
suggestions.append({'disk': disk, 'by_id': alias,
'vgs': candidate['vgs']})
except (ValueError, RuntimeError) as exc:
blocked.append({'disk': disk, 'reason': str(exc)})
elif re.fullmatch(r'/dev/vd[a-z]+', disk):
try:
candidate = discover([disk], include_suggestions=False)
suggestions.append({'disk': disk, 'by_id': disk,
'vgs': candidate['vgs']})
except (ValueError, RuntimeError) as exc:
blocked.append({'disk': disk, 'reason': str(exc)})
else:
blocked.append({'disk': disk, 'reason': '缺少可用的 /dev/disk/by-id/ 整盘稳定路径'})
plan['suggested_targets'] = suggestions
plan['blocked_targets'] = blocked
digest = hashlib.sha256(json.dumps(plan, sort_keys=True).encode()).hexdigest()
plan['fingerprint'] = digest
return plan
def execute(plan):
completed = []
for vg in plan['vgs']:
for lv in plan['lvs'][vg]:
command(['lvremove', '-ff', '-y', lv])
completed.append('lvremove ' + lv)
command(['vgremove', '-ff', '-y', vg])
completed.append('vgremove ' + vg)
for pv in plan['pvs'][vg]:
command(['pvremove', '-ff', '-y', pv])
completed.append('pvremove ' + pv)
for target in plan['disks']:
command(['wipefs', '-af', target['requested']])
completed.append('wipefs -af ' + target['requested'])
return completed
if __name__ == '__main__':
try:
mode = sys.argv[1]
requested = json.loads(sys.argv[2])
if mode not in ('plan', 'execute'):
raise ValueError('未知模式')
plan = discover(requested)
if mode == 'plan':
print(json.dumps(plan, ensure_ascii=False))
else:
if len(sys.argv) != 5 or sys.argv[4] != 'ERASE_LVM_AND_WIPEFS':
raise ValueError('缺少明确确认字符串')
if not requested or plan['fingerprint'] != sys.argv[3]:
raise ValueError('设备状态已变化或未选择目标;拒绝执行')
print(json.dumps({'completed': execute(plan)}, ensure_ascii=False))
except (ValueError, RuntimeError, OSError, subprocess.TimeoutExpired) as exc:
print(str(exc), file=sys.stderr)
sys.exit(1)
5. 验证与已知边界
生成笔记前,两个 Playbook 均已通过 ansible-playbook --syntax-check,两个 Python 文件均已通过 python3 -m py_compile。LVM role 的只读发现曾在用户提供的五台主机上成功运行;最新的 VirtIO /dev/vdX 支持和 always 清理逻辑尚未在这些主机上复测。健康巡检 role 尚未在目标 Linux 主机上实测。版本不同的 lsblk、LVM2 或设备映射可能使发现失败;失败时不会进入 LVM 删除步骤。