Ansible Roles:Linux 健康巡检与 LVM 磁盘清理

Ansible Roles:Linux 健康巡检与 LVM 磁盘清理

更新日期:2026-09-28
源代码目录:/Users/qiaoning5/PycharmProjects/PythonProject

这份笔记保存了本次编写的两个 role 的完整源码与运行方法。下方源码是生成笔记时的快照;以后若修改项目文件,应重新同步笔记。

1. 文件结构

PythonProject/
├── health-inspector.yml
├── lvm-disk-cleanup.yml
└── roles/
    ├── health-inspector/
    │   ├── defaults/main.yml
    │   ├── tasks/main.yml
    │   └── files/collect.py
    └── lvm-disk-cleanup/
        ├── defaults/main.yml
        ├── tasks/main.yml
        └── files/cleanup.py

两个 Playbook 都使用 become: true。目标机需要 Linux 和 Python 3;控制机需要 Ansible。

2. Linux 健康巡检 role

覆盖指标

类别 指标
硬件 硬盘 SMART 5/187/197/198、温度、ECC CE/UCE、电源和风扇传感器
系统资源 1/5/15 分钟负载与每 CPU 负载、运行队列、I/O wait、MemAvailable、Swap、si/so、OOM、磁盘空间和 inode、iostat 的 util/await/IOPS/队列长度
网络 ESTABLISHED、TIME_WAIT、CLOSE_WAIT、网卡错误与丢包、指定端口监听
服务 指定服务 active/enabled、systemctl failed units
安全 last/lastb、认证失败日志、UID 0 和空密码账户、口令期限、SSH root/密码认证、SUID/SGID 列表
日志 dmesg error/fail/warn、本次启动 journal error、core dump 文件

每项在 JSON 中有 value 与 status。状态为 ok、warning、critical 或 unknown。SMART、IPMI、iostat 等依赖可选工具;缺少工具或硬件接口会标记为 unknown,不会误报正常。CE/UCE、网卡错误等是累计计数;iostat 使用 1 秒采样值。

运行

cd /Users/qiaoning5/PycharmProjects/PythonProject
ansible-playbook -i inventory health-inspector.yml \
  -e '{"health_inspector_services":["sshd","nginx"],"health_inspector_ports":[22,80,443]}'

配置在 roles/health-inspector/defaults/main.yml。可用 -e 覆盖服务、端口、阈值、扫描范围和报告目录。报告默认保存在控制端 health_reports/<主机名>.json,权限为 0600。目标机安装 smartmontools、ipmitool、sysstat 后可获取更完整的数据。报告含登录记录和账户信息,应限制访问。

3. LVM 磁盘清理 role

工作流程

  1. 使用 lsblk、pvs、lvs 识别物理磁盘、PV、VG、LV 的关系。
  2. 默认只显示计划与可执行命令,不删除磁盘数据。
  3. 执行模式先核对目标整盘、挂载点、Swap、VG 跨盘情况和确认字符串。
  4. 执行前再次发现并比较计划指纹;发现变化则拒绝继续。
  5. 依次运行 lvremove -ff -y、vgremove -ff -y、pvremove -ff -y、wipefs -af。
  6. always 任务删除目标机 /tmp/ansible-lvm-disk-cleanup.py。如目标机断线,Ansible 无法完成清理,恢复连接后可重新运行默认发现命令或单独删除该文件。

这项操作不可逆。删除 docker VG 会删除其全部 Docker LV 和数据。wipefs 清除磁盘签名,并不保证安全擦除磁盘上所有数据块。role 接受 /dev/disk/by-id/ 整盘路径;无稳定路径的 VirtIO 整盘可明确使用 /dev/vdX。它拒绝分区路径、已挂载磁盘、Swap、跨越未选择磁盘的 VG,以及执行前发生变化的计划。

默认只读发现

你在有 ip.txt 的目录运行:

ansible-playbook -i ip.txt /Users/qiaoning5/PycharmProjects/PythonProject/lvm-disk-cleanup.yml

默认值是 lvm_cleanup_disks: []、lvm_cleanup_execute: false、lvm_cleanup_confirm: ''。默认运行会上传临时脚本进行发现,最后删除临时脚本;不会调用删除命令。输出中的 inventory 显示 VG/PV/LV/磁盘关系,suggested_targets 表示通过安全检查的目标,blocked_targets 说明不能生成命令的原因。

执行单台主机的 /dev/vdb

仅在确认该主机的 /dev/vdb 是要清空的整盘、业务已停止并有备份后运行。以下命令只针对 11.50.138.135;其他主机替换 --limit 值,逐台执行与核对结果。

ansible-playbook -i ip.txt /Users/qiaoning5/PycharmProjects/PythonProject/lvm-disk-cleanup.yml \
  --limit 11.50.138.135 \
  -e '{"lvm_cleanup_disks":["/dev/vdb"],"lvm_cleanup_execute":true,"lvm_cleanup_confirm":"ERASE_LVM_AND_WIPEFS"}'

本次用户贴出的只读发现结果显示:11.50.138.134、11.50.138.135、11.50.139.1、11.50.139.122 的 /dev/vdb 属于 docker VG;11.50.139.26 没有发现 PV。该结果是历史快照,执行前仍以目标机当前发现结果为准。若使用了另一份 disk/lvm-disk-cleanup.yml,需确认其 role 源码已同步本笔记中的最新版本。

单独删除旧运行遗留的临时文件

ansible -i ip.txt all -b -m file -a 'path=/tmp/ansible-lvm-disk-cleanup.py state=absent'

这条命令只删除临时 Python 文件,不删除 LVM 或磁盘数据。

4. 完整源码快照

以下代码块按实际目录组织,复制到同名文件即可还原本次两个 role。

health-inspector.yml

---
- name: Linux 健康巡检
  hosts: all
  become: true
  gather_facts: false
  roles:
    - health-inspector

roles/health-inspector/defaults/main.yml

---
health_inspector_report_dir: "{{ playbook_dir }}/health_reports"
health_inspector_services: []
health_inspector_ports: []
health_inspector_suid_paths: ['/usr/bin', '/usr/sbin', '/bin', '/sbin']
health_inspector_core_paths: ['/var/lib/systemd/coredump', '/var/crash', '/var/core']
health_inspector_log_lines: 30
health_inspector_thresholds:
  smart_temperature_warn_c: 50
  smart_temperature_critical_c: 60
  load_per_cpu_warn: 0.7
  load_per_cpu_critical: 1.0
  iowait_warn_percent: 20
  memory_available_warn_percent: 10
  swap_used_warn_percent: 50
  filesystem_warn_percent: 80
  filesystem_critical_percent: 95
  inode_warn_percent: 80
  disk_util_warn_percent: 80
  disk_await_warn_ms: 50
  time_wait_warn: 5000
  close_wait_warn: 100

roles/health-inspector/tasks/main.yml

---
- name: 创建控制端报告目录
  ansible.builtin.file:
    path: "{{ health_inspector_report_dir }}"
    state: directory
    mode: '0700'
  delegate_to: localhost
  become: false

- name: 上传只读巡检程序
  ansible.builtin.copy:
    src: collect.py
    dest: /tmp/ansible-health-inspector-collect.py
    mode: '0700'

- name: 采集健康指标
  ansible.builtin.command:
    argv:
      - python3
      - /tmp/ansible-health-inspector-collect.py
      - "{{ health_inspector_services | to_json }}"
      - "{{ health_inspector_ports | to_json }}"
      - "{{ health_inspector_suid_paths | to_json }}"
      - "{{ health_inspector_core_paths | to_json }}"
      - "{{ health_inspector_thresholds | to_json }}"
      - "{{ health_inspector_log_lines | int | string }}"
  register: health_inspector_raw
  changed_when: false
  failed_when: health_inspector_raw.rc != 0

- name: 解析巡检结果
  ansible.builtin.set_fact:
    health_inspector_result: "{{ health_inspector_raw.stdout | from_json }}"

- name: 保存本地 JSON 报告
  ansible.builtin.copy:
    content: "{{ health_inspector_result | to_nice_json }}\n"
    dest: "{{ health_inspector_report_dir }}/{{ inventory_hostname | regex_replace('[^A-Za-z0-9_.-]', '_') }}.json"
    mode: '0600'
  delegate_to: localhost
  become: false

- name: 显示巡检汇总
  ansible.builtin.debug:
    msg:
      host: "{{ inventory_hostname }}"
      status: "{{ health_inspector_result.status }}"
      critical: "{{ health_inspector_result.summary.critical }}"
      warning: "{{ health_inspector_result.summary.warning }}"
      unknown: "{{ health_inspector_result.summary.unknown }}"
      report: "{{ health_inspector_report_dir }}/{{ inventory_hostname | regex_replace('[^A-Za-z0-9_.-]', '_') }}.json"

- name: 删除目标机临时巡检程序
  ansible.builtin.file:
    path: /tmp/ansible-health-inspector-collect.py
    state: absent
  changed_when: false

roles/health-inspector/files/collect.py

#!/usr/bin/env python3
"""Linux read-only health snapshot. Optional tools yield unknown, never healthy."""
import datetime
import json
import os
import re
import shutil
import subprocess
import sys
from pathlib import Path


def run(args, timeout=12):
    if not shutil.which(args[0]):
        return None
    try:
        p = subprocess.run(args, text=True, stdout=subprocess.PIPE,
                           stderr=subprocess.PIPE, timeout=timeout, check=False)
        return p.stdout if p.returncode == 0 else None
    except (OSError, subprocess.TimeoutExpired):
        return None


def read(path):
    try:
        return Path(path).read_text(errors='replace')
    except OSError:
        return None


def number(value):
    try:
        return float(value)
    except (TypeError, ValueError):
        return None


def sample(text, limit):
    return (text or '').splitlines()[-limit:]


def add(group, name, value, status='ok', detail=None):
    item = {'value': value, 'status': status}
    if detail is not None:
        item['detail'] = detail
    result['metrics'].setdefault(group, {})[name] = item


def graded(value, warn=None, critical=None, reverse=False):
    if value is None:
        return 'unknown'
    if reverse:
        return 'critical' if critical is not None and value <= critical else ('warning' if warn is not None and value <= warn else 'ok')
    return 'critical' if critical is not None and value >= critical else ('warning' if warn is not None and value >= warn else 'ok')


def hardware():
    devices = []
    raw = run(['lsblk', '-J', '-d', '-o', 'NAME,TYPE,TRAN'])
    if raw:
        try:
            devices = ['/dev/' + d['name'] for d in json.loads(raw)['blockdevices'] if d.get('type') == 'disk']
        except (KeyError, ValueError, TypeError):
            pass
    for dev in devices:
        # smartctl uses nonzero bitmask exit codes even when valid SMART data exists.
        try:
            raw = subprocess.run(['smartctl', '-j', '-a', dev], text=True,
                                 capture_output=True, timeout=25).stdout if shutil.which('smartctl') else None
        except (OSError, subprocess.TimeoutExpired):
            raw = None
        if raw is None:
            add('hardware', dev + '.smart', None, 'unknown', 'smartctl 缺失、设备不支持或访问失败')
            continue
        try:
            data = json.loads(raw)
        except ValueError:
            add('hardware', dev + '.smart', None, 'unknown', 'SMART JSON 无效')
            continue
        table = {x.get('id'): x for x in data.get('ata_smart_attributes', {}).get('table', [])}
        for ident, label in [(5, 'Reallocated_Sector_Ct'), (187, 'Reported_Uncorrect'),
                             (197, 'Current_Pending_Sector'), (198, 'Offline_Uncorrectable')]:
            v = number(table.get(ident, {}).get('raw', {}).get('value'))
            add('hardware', dev + '.smart_' + str(ident) + '_' + label, v,
                graded(v, 1, 1), '原始计数;NVMe 或不支持的设备为 unknown')
        temp = number(data.get('temperature', {}).get('current'))
        add('hardware', dev + '.temperature_c', temp,
            graded(temp, th['smart_temperature_warn_c'], th['smart_temperature_critical_c']))
    if not devices:
        add('hardware', 'smart_devices', None, 'unknown', 'lsblk 未发现磁盘或不可用')

    edac = Path('/sys/devices/system/edac/mc')
    for suffix, label in [('ce_count', 'ecc_ce'), ('ue_count', 'ecc_uce')]:
        vals = [number(read(p).strip()) for p in edac.glob('mc[0-9]*/' + suffix)] if edac.exists() else []
        total = int(sum(vals)) if vals and all(v is not None for v in vals) else None
        add('hardware', label, total, graded(total, 1, 1), '开机以来累计次数')

    raw = run(['ipmitool', 'sdr', 'elist', 'all'], 20)
    if raw is None:
        add('hardware', 'power_supplies', None, 'unknown', '需要 IPMI/BMC 和 ipmitool')
        add('hardware', 'fans', None, 'unknown', '需要 IPMI/BMC 和 ipmitool')
    else:
        for key, regex in [('power_supplies', r'(?i)\b(?:psu|power supply|pwr)\b'),
                           ('fans', r'(?i)\bfan\b')]:
            lines = [x for x in raw.splitlines() if re.search(regex, x)]
            bad = [x for x in lines if re.search(r'(?i)\b(?:fail|critical|not present|no reading|nr|cr|nc)\b', x)]
            incomplete_psu = key == 'power_supplies' and len(lines) < 2
            add('hardware', key, lines or None, 'unknown' if not lines or incomplete_psu else ('critical' if bad else 'ok'),
                '传感器原始状态;双路电源需核对两条 PSU 记录')


def resources():
    load = os.getloadavg()
    cpu = os.cpu_count() or 1
    for index, interval in enumerate((1, 5, 15)):
        add('resources', 'load_' + str(interval) + 'm', round(load[index], 2))
        ratio = round(load[index] / cpu, 3)
        add('resources', 'load_' + str(interval) + 'm_per_cpu', ratio,
            graded(ratio, th['load_per_cpu_warn'], th['load_per_cpu_critical']))
    raw = run(['vmstat', '1', '2'])
    if raw:
        rows = raw.splitlines()
        cols = rows[-1].split() if len(rows) >= 4 else []
        # Linux vmstat: r b swpd free buff cache si so bi bo in cs us sy id wa st
        for key, pos, warn in [('run_queue', 0, cpu), ('swap_in_kb_s', 6, 1),
                                ('swap_out_kb_s', 7, 1), ('cpu_iowait_percent', 15, th['iowait_warn_percent'])]:
            v = number(cols[pos]) if len(cols) > pos else None
            add('resources', key, v, graded(v, warn))
    else:
        for key in ('run_queue', 'swap_in_kb_s', 'swap_out_kb_s', 'cpu_iowait_percent'):
            add('resources', key, None, 'unknown', 'vmstat 不可用')
    mem = {m.group(1): int(m.group(2)) for m in re.finditer(r'^(\w+):\s+(\d+) kB', read('/proc/meminfo') or '', re.M)}
    available = mem.get('MemAvailable')
    pct = round(available / mem['MemTotal'] * 100, 2) if available is not None and mem.get('MemTotal') else None
    add('resources', 'memory_available_kb', available, 'unknown' if available is None else 'ok')
    add('resources', 'memory_available_percent', pct, graded(pct, th['memory_available_warn_percent'], reverse=True))
    swap = mem.get('SwapTotal', 0)
    used = swap - mem.get('SwapFree', 0)
    add('resources', 'swap_used_kb', used)
    add('resources', 'swap_used_percent', round(used / swap * 100, 2) if swap else 0,
        graded(used / swap * 100 if swap else 0, th['swap_used_warn_percent']))
    journal = run(['journalctl', '-k', '-b', '--no-pager', '-q'], 20)
    oom = [x for x in (journal or '').splitlines() if re.search(r'(?i)oom-killer|out of memory|killed process', x)]
    add('resources', 'oom_events_current_boot', sample('\n'.join(oom), limit) if journal is not None else None,
        'unknown' if journal is None else graded(len(oom), 1, 1))

    for mode, key, threshold in [([], 'filesystem', 'filesystem_warn_percent'),
                                 (['-i'], 'inodes', 'inode_warn_percent')]:
        raw = run(['df', '-P'] + mode)
        if raw is None:
            add('resources', key, None, 'unknown', 'df 不可用')
            continue
        items = []
        for line in raw.splitlines()[1:]:
            parts = line.split()
            if len(parts) >= 6 and parts[4].endswith('%') and not parts[0].startswith(('tmpfs', 'devtmpfs')):
                items.append({'mount': parts[-1], 'used_percent': int(parts[4][:-1])})
        status = 'critical' if key == 'filesystem' and any(x['used_percent'] >= th['filesystem_critical_percent'] for x in items) else graded(max((x['used_percent'] for x in items), default=None), th[threshold])
        add('resources', key, items, status)

    raw = run(['iostat', '-dx', '-y', '1', '2'], 15)
    if raw is None:
        add('resources', 'disk_io', None, 'unknown', '需要 sysstat/iostat')
    else:
        lines = raw.splitlines()
        headers = [i for i, x in enumerate(lines) if x.strip().startswith('Device')]
        items = []
        if headers:
            header = lines[headers[-1]].split()
            for line in lines[headers[-1] + 1:]:
                cols = line.split()
                if len(cols) != len(header):
                    continue
                row = dict(zip(header, cols))
                def field(*names):
                    return next((number(row[n]) for n in names if n in row), None)
                rs, ws = field('r/s'), field('w/s')
                item = {'device': cols[0], 'util_percent': field('%util'),
                        'await_ms': field('await'), 'iops': rs + ws if rs is not None and ws is not None else None,
                        'avgqu_sz': field('avgqu-sz', 'aqu-sz')}
                items.append(item)
        warn = any((x['util_percent'] or 0) >= th['disk_util_warn_percent'] or
                   (x['await_ms'] or 0) >= th['disk_await_warn_ms'] for x in items)
        add('resources', 'disk_io', items or None, 'unknown' if not items else ('warning' if warn else 'ok'),
            '采样区间 1 秒;iostat -y 排除开机以来平均值')


def network():
    raw = run(['ss', '-tan'])
    if raw is None:
        add('network', 'tcp_states', None, 'unknown', 'ss 不可用')
    else:
        counts = {state: len(re.findall(r'^' + state + r'\s', raw, re.M)) for state in ('ESTAB', 'TIME-WAIT', 'CLOSE-WAIT')}
        add('network', 'established', counts['ESTAB'])
        add('network', 'time_wait', counts['TIME-WAIT'], graded(counts['TIME-WAIT'], th['time_wait_warn']))
        add('network', 'close_wait', counts['CLOSE-WAIT'], graded(counts['CLOSE-WAIT'], th['close_wait_warn']))
    nic = {}
    for path in Path('/sys/class/net').glob('*'):
        if path.name == 'lo':
            continue
        nic[path.name] = {}
        for field in ('rx_errors', 'tx_errors', 'rx_dropped', 'tx_dropped'):
            raw = read(path / 'statistics' / field)
            nic[path.name][field] = int(raw.strip()) if raw and raw.strip().isdigit() else None
    add('network', 'interfaces', nic or None,
        'unknown' if not nic else ('warning' if any((v or 0) > 0 for fields in nic.values() for v in fields.values()) else 'ok'),
        '开机以来累计计数')
    listening = run(['ss', '-ltnH'])
    for port in ports:
        found = listening is not None and any(re.search(r':' + str(port) + r'\s', line) for line in listening.splitlines())
        add('network', 'listen_' + str(port), found if listening is not None else None,
            'unknown' if listening is None else ('ok' if found else 'critical'))


def services():
    for svc in svc_names:
        # systemctl returns nonzero for inactive/disabled; query output remains useful.
        for key, command, expected in [('active', 'is-active', 'active'), ('enabled', 'is-enabled', 'enabled')]:
            try:
                p = subprocess.run(['systemctl', command, svc], capture_output=True, text=True, timeout=8)
                val = p.stdout.strip() or 'unknown'
            except (OSError, subprocess.TimeoutExpired):
                val = 'unknown'
            add('services', svc + '.' + key, val, 'unknown' if val == 'unknown' else ('ok' if val == expected else 'critical'))
    raw = run(['systemctl', '--failed', '--no-legend', '--plain'])
    failed = [x.strip() for x in (raw or '').splitlines() if x.strip() and not x.startswith('0 loaded')]
    add('services', 'failed_units', failed if raw is not None else None,
        'unknown' if raw is None else graded(len(failed), 1, 1))


def security():
    for key, command in [('successful_logins', ['last', '-n', str(limit)]),
                         ('failed_logins', ['lastb', '-n', str(limit)])]:
        raw = run(command)
        add('security', key, sample(raw, limit) if raw is not None else None,
            'unknown' if raw is None else 'ok')
    auth = next((read(p) for p in ('/var/log/secure', '/var/log/auth.log') if read(p) is not None), None)
    failures = [x for x in (auth or '').splitlines() if re.search(r'Failed password|authentication failure', x, re.I)]
    add('security', 'auth_failures', sample('\n'.join(failures), limit) if auth is not None else None,
        'unknown' if auth is None else graded(len(failures), 1))
    passwd = read('/etc/passwd')
    shadow = read('/etc/shadow')
    uid0 = [x.split(':')[0] for x in (passwd or '').splitlines() if len(x.split(':')) > 2 and x.split(':')[2] == '0']
    empty = [x.split(':')[0] for x in (shadow or '').splitlines() if len(x.split(':')) > 1 and x.split(':')[1] == '']
    add('security', 'uid0_accounts', uid0 if passwd is not None else None,
        'unknown' if passwd is None else ('warning' if len(uid0) > 1 else 'ok'))
    add('security', 'empty_password_accounts', empty if shadow is not None else None,
        'unknown' if shadow is None else graded(len(empty), 1, 1))
    defs = read('/etc/login.defs')
    policy = {}
    for key in ('PASS_MAX_DAYS', 'PASS_MIN_DAYS', 'PASS_WARN_AGE'):
        m = re.search(r'^\s*' + key + r'\s+(\S+)', defs or '', re.M)
        policy[key] = m.group(1) if m else None
    add('security', 'password_age_policy', policy, 'unknown' if defs is None else 'ok',
        'login.defs 默认值;已有账户策略可用 chage -l 单独核查')
    ssh = run(['sshd', '-T']) or run(['/usr/sbin/sshd', '-T'])
    for key, expected in [('permitrootlogin', 'no'), ('passwordauthentication', 'no')]:
        m = re.search(r'^' + key + r'\s+(\S+)', ssh or '', re.M | re.I)
        value = m.group(1) if m else None
        add('security', 'ssh_' + key, value, 'unknown' if value is None else ('ok' if value == expected else 'warning'))
    suid = []
    for base in suid_paths:
        if not os.path.isdir(base):
            continue
        for root, dirs, files in os.walk(base):
            dirs[:] = [d for d in dirs if not os.path.islink(os.path.join(root, d))]
            for name in files:
                path = os.path.join(root, name)
                try:
                    mode = os.stat(path, follow_symlinks=False).st_mode
                    if mode & 0o6000:
                        suid.append({'path': path, 'mode': oct(mode & 0o7777)})
                except OSError:
                    pass
    add('security', 'suid_sgid_files', suid, 'ok', '用于基线对比;存在 SUID/SGID 本身不判异常')


def logs():
    raw = run(['dmesg', '--color=never'])
    hits = [x for x in (raw or '').splitlines() if re.search(r'\b(error|fail|warn)\b', x, re.I)]
    add('logs', 'dmesg_error_fail_warn', sample('\n'.join(hits), limit) if raw is not None else None,
        'unknown' if raw is None else graded(len(hits), 1))
    raw = run(['journalctl', '-b', '-p', 'err', '--no-pager', '-q', '-n', str(limit)], 20)
    lines = [x for x in (raw or '').splitlines() if x.strip()]
    add('logs', 'journal_current_boot_errors', lines if raw is not None else None,
        'unknown' if raw is None else graded(len(lines), 1))
    cores = []
    for base in core_paths:
        if os.path.isdir(base):
            for root, dirs, files in os.walk(base):
                for name in files:
                    path = os.path.join(root, name)
                    try:
                        st = os.stat(path)
                        cores.append({'path': path, 'size_bytes': st.st_size})
                    except OSError:
                        pass
    add('logs', 'core_dump_files', cores, graded(len(cores), 1))


if __name__ == '__main__':
    svc_names, ports, suid_paths, core_paths, th = [json.loads(x) for x in sys.argv[1:6]]
    limit = max(1, min(200, int(sys.argv[6])))
    result = {'host': os.uname().nodename, 'collected_at': datetime.datetime.now(datetime.timezone.utc).isoformat(), 'metrics': {}}
    for collector in (hardware, resources, network, services, security, logs):
        try:
            collector()
        except Exception as exc:
            add('collection_errors', collector.__name__, str(exc), 'unknown')
    summary = {s: sum(x['status'] == s for group in result['metrics'].values() for x in group.values())
               for s in ('ok', 'warning', 'critical', 'unknown')}
    result['summary'] = summary
    result['status'] = 'critical' if summary['critical'] else ('warning' if summary['warning'] else ('unknown' if summary['unknown'] else 'ok'))
    print(json.dumps(result, ensure_ascii=False))

lvm-disk-cleanup.yml

---
- name: 清理明确指定的 LVM 物理磁盘
  hosts: all
  become: true
  gather_facts: false
  roles:
    - lvm-disk-cleanup

roles/lvm-disk-cleanup/defaults/main.yml

---
# 优先使用 /dev/disk/by-id;无稳定路径的 VirtIO 整盘可显式指定 /dev/vdX。
lvm_cleanup_disks: []
lvm_cleanup_confirm: ''
lvm_cleanup_execute: false
# 默认沿用本次运行的 inventory;也可显式覆盖。
lvm_cleanup_inventory_path: "{{ ansible_inventory_sources[0] }}"

roles/lvm-disk-cleanup/tasks/main.yml

---
- name: 运行 LVM 巡检与清理并复原临时环境
  block:
    - name: 上传 LVM 发现与清理程序
      ansible.builtin.copy:
        src: cleanup.py
        dest: /tmp/ansible-lvm-disk-cleanup.py
        mode: '0700'

    - name: 发现 LVM 磁盘并执行安全检查
      ansible.builtin.command:
        argv:
          - python3
          - /tmp/ansible-lvm-disk-cleanup.py
          - plan
          - "{{ lvm_cleanup_disks | to_json }}"
      register: lvm_cleanup_plan_raw
      changed_when: false

    - name: 导出清理计划
      ansible.builtin.set_fact:
        lvm_cleanup_plan: "{{ lvm_cleanup_plan_raw.stdout | from_json }}"

    - name: 显示清理计划
      ansible.builtin.debug:
        var: lvm_cleanup_plan

    - name: 输出已通过安全检查的完整执行命令
      ansible.builtin.debug:
        msg: >-
          ansible-playbook -i {{ lvm_cleanup_inventory_path | quote }} {{ (playbook_dir ~ '/lvm-disk-cleanup.yml') | quote }}
          --limit {{ inventory_hostname | quote }}
          -e {{ {'lvm_cleanup_disks': [item.by_id], 'lvm_cleanup_execute': true,
                 'lvm_cleanup_confirm': 'ERASE_LVM_AND_WIPEFS'} | to_json | quote }}
      loop: "{{ lvm_cleanup_plan.suggested_targets | default([]) }}"
      loop_control:
        label: "{{ item.disk }}"
      when: not (lvm_cleanup_execute | bool)

    - name: 提示无可直接清理的磁盘
      ansible.builtin.debug:
        msg: '未发现可生成执行命令的磁盘;可能有挂载、Swap、VG 跨盘或缺少 /dev/disk/by-id 路径。'
      when:
        - not (lvm_cleanup_execute | bool)
        - lvm_cleanup_plan.suggested_targets | default([]) | length == 0

    - name: 显示不能生成执行命令的具体原因
      ansible.builtin.debug:
        msg: "{{ item.disk }}: {{ item.reason }}"
      loop: "{{ lvm_cleanup_plan.blocked_targets | default([]) }}"
      loop_control:
        label: "{{ item.disk }}"
      when: not (lvm_cleanup_execute | bool)

    - name: 仅在明确授权时执行清理
      when: lvm_cleanup_execute | bool
      block:
        - name: 核对清理授权
          ansible.builtin.assert:
            that:
              - lvm_cleanup_confirm == 'ERASE_LVM_AND_WIPEFS'
              - lvm_cleanup_plan.disks | length > 0
            fail_msg: '必须明确设置目标盘、lvm_cleanup_execute=true 和确认字符串。'

        - name: 删除 LV/VG/PV 并擦除磁盘签名
          ansible.builtin.command:
            argv:
              - python3
              - /tmp/ansible-lvm-disk-cleanup.py
              - execute
              - "{{ lvm_cleanup_disks | to_json }}"
              - "{{ lvm_cleanup_plan.fingerprint }}"
              - "{{ lvm_cleanup_confirm }}"
          register: lvm_cleanup_result
          changed_when: true

        - name: 显示执行结果
          ansible.builtin.debug:
            msg: "{{ lvm_cleanup_result.stdout | from_json }}"
  always:
    - name: 删除目标机临时巡检程序
      ansible.builtin.file:
        path: /tmp/ansible-lvm-disk-cleanup.py
        state: absent
      changed_when: false

roles/lvm-disk-cleanup/files/cleanup.py

#!/usr/bin/env python3
"""Discover LVM disk ownership; destroy only explicitly selected, idle disks."""
import hashlib
import json
import os
import re
import shutil
import subprocess
import sys


def command(argv):
    proc = subprocess.run(argv, text=True, stdout=subprocess.PIPE,
                          stderr=subprocess.PIPE, timeout=60, check=False)
    if proc.returncode:
        raise RuntimeError('%s: %s' % (' '.join(argv), proc.stderr.strip()))
    return proc.stdout


def report(command_name, columns):
    data = json.loads(command([command_name, '--reportformat', 'json',
                               '--units', 'b', '--nosuffix', '-o', columns]))
    return data['report'][0][{'pvs': 'pv', 'lvs': 'lv'}[command_name]]


def discover(requested, include_suggestions=True):
    if not isinstance(requested, list) or any(not isinstance(x, str) for x in requested):
        raise ValueError('lvm_cleanup_disks 必须是设备路径列表')
    if len(requested) != len(set(requested)):
        raise ValueError('目标磁盘重复')
    for tool in ('lsblk', 'pvs', 'lvs', 'lvremove', 'vgremove', 'pvremove', 'wipefs'):
        if shutil.which(tool) is None:
            raise RuntimeError('缺少命令: ' + tool)
    # MOUNTPOINTS (plural), SERIAL and WWN are absent on older util-linux.
    tree = json.loads(command(['lsblk', '-J', '-p', '-o', 'NAME,TYPE,MOUNTPOINT']))['blockdevices']
    by_path = {}
    disks = {}

    def walk(node, disk):
        path = os.path.realpath(node['name'])
        if node['type'] == 'disk':
            disk = path
            disks[disk] = {'path': path, 'nodes': []}
        if disk is not None:
            mountpoint = node.get('mountpoint')
            mountpoints = [mountpoint] if mountpoint else []
            disks[disk]['nodes'].append({'path': path, 'type': node['type'],
                                         'mountpoints': mountpoints})
            by_path[path] = disk
        for child in node.get('children') or []:
            walk(child, disk)

    for root in tree:
        walk(root, None)

    pvs = report('pvs', 'pv_name,vg_name')
    lvs = report('lvs', 'lv_path,vg_name')
    pv_by_vg = {}
    for pv in pvs:
        vg = (pv.get('vg_name') or '').strip()
        if vg:
            pv_by_vg.setdefault(vg, []).append(os.path.realpath(pv['pv_name'].strip()))
    lv_by_vg = {}
    for lv in lvs:
        vg = (lv.get('vg_name') or '').strip()
        path = (lv.get('lv_path') or '').strip()
        if vg and path:
            lv_by_vg.setdefault(vg, []).append(path)

    chosen = []
    for name in requested:
        stable = name.startswith('/dev/disk/by-id/')
        virtio = re.fullmatch(r'/dev/vd[a-z]+', name) is not None
        if not stable and not virtio:
            raise ValueError('目标必须使用 /dev/disk/by-id/ 或明确的 /dev/vdX 整盘路径: ' + name)
        if not os.path.exists(name):
            raise ValueError('设备不存在: ' + name)
        path = os.path.realpath(name)
        if path not in disks:
            raise ValueError('目标不是整块物理磁盘: ' + name)
        if virtio and path not in {pv for paths in pv_by_vg.values() for pv in paths}:
            raise ValueError('VirtIO 整盘自身不是已加入 VG 的 PV: ' + name)
        if path in chosen:
            raise ValueError('不同路径解析到同一磁盘: ' + path)
        chosen.append(path)

    swap_paths = set()
    with open('/proc/swaps', encoding='utf-8') as stream:
        for line in stream.readlines()[1:]:
            swap_paths.add(os.path.realpath(line.split()[0]))

    all_pv_disks = {}
    for vg, paths in pv_by_vg.items():
        all_pv_disks[vg] = [by_path.get(p) for p in paths]
    targets = []
    for disk in chosen:
        nodes = disks[disk]['nodes']
        if any(n['mountpoints'] for n in nodes):
            raise ValueError('目标盘有挂载点: ' + disk)
        if any(n['path'] in swap_paths for n in nodes):
            raise ValueError('目标盘正在用于 Swap: ' + disk)
        related = sorted(vg for vg, owners in all_pv_disks.items() if disk in owners)
        if not related:
            raise ValueError('目标盘没有已加入 VG 的 PV: ' + disk)
        targets.append({'requested': requested[chosen.index(disk)], 'disk': disk,
                        'vgs': related})
    target_vgs = sorted({vg for target in targets for vg in target['vgs']})
    for vg in target_vgs:
        owners = all_pv_disks[vg]
        if any(owner is None or owner not in chosen for owner in owners):
            raise ValueError('VG %s 跨越未选择或无法识别的磁盘;必须整体选择' % vg)
        if any(os.path.realpath(lv) in swap_paths for lv in lv_by_vg.get(vg, [])):
            raise ValueError('VG %s 中有正在使用的 Swap LV' % vg)
    inventory = [{'vg': vg, 'pvs': pv_by_vg[vg], 'disks': all_pv_disks[vg],
                  'lvs': sorted(lv_by_vg.get(vg, []))} for vg in sorted(pv_by_vg)]
    plan = {'disks': targets, 'vgs': target_vgs, 'inventory': inventory,
            'lvs': {vg: sorted(lv_by_vg.get(vg, [])) for vg in target_vgs},
            'pvs': {vg: sorted(pv_by_vg[vg]) for vg in target_vgs},
            'discovered_pv_disks': sorted(set(x for owners in all_pv_disks.values() for x in owners if x))}
    if include_suggestions and not requested:
        suggestions = []
        blocked = []
        # Only stable whole-disk aliases, never partition aliases.
        alias_dir = '/dev/disk/by-id'
        aliases = {}
        if os.path.isdir(alias_dir):
            for name in sorted(os.listdir(alias_dir)):
                if '-part' in name or not name.startswith(('wwn-', 'nvme-eui.', 'ata-', 'scsi-')):
                    continue
                alias = os.path.join(alias_dir, name)
                disk = os.path.realpath(alias)
                if disk in disks and disk not in aliases:
                    aliases[disk] = alias
        for disk in plan['discovered_pv_disks']:
            alias = aliases.get(disk)
            if alias:
                try:
                    candidate = discover([alias], include_suggestions=False)
                    suggestions.append({'disk': disk, 'by_id': alias,
                                        'vgs': candidate['vgs']})
                except (ValueError, RuntimeError) as exc:
                    blocked.append({'disk': disk, 'reason': str(exc)})
            elif re.fullmatch(r'/dev/vd[a-z]+', disk):
                try:
                    candidate = discover([disk], include_suggestions=False)
                    suggestions.append({'disk': disk, 'by_id': disk,
                                        'vgs': candidate['vgs']})
                except (ValueError, RuntimeError) as exc:
                    blocked.append({'disk': disk, 'reason': str(exc)})
            else:
                blocked.append({'disk': disk, 'reason': '缺少可用的 /dev/disk/by-id/ 整盘稳定路径'})
        plan['suggested_targets'] = suggestions
        plan['blocked_targets'] = blocked
    digest = hashlib.sha256(json.dumps(plan, sort_keys=True).encode()).hexdigest()
    plan['fingerprint'] = digest
    return plan


def execute(plan):
    completed = []
    for vg in plan['vgs']:
        for lv in plan['lvs'][vg]:
            command(['lvremove', '-ff', '-y', lv])
            completed.append('lvremove ' + lv)
        command(['vgremove', '-ff', '-y', vg])
        completed.append('vgremove ' + vg)
        for pv in plan['pvs'][vg]:
            command(['pvremove', '-ff', '-y', pv])
            completed.append('pvremove ' + pv)
    for target in plan['disks']:
        command(['wipefs', '-af', target['requested']])
        completed.append('wipefs -af ' + target['requested'])
    return completed


if __name__ == '__main__':
    try:
        mode = sys.argv[1]
        requested = json.loads(sys.argv[2])
        if mode not in ('plan', 'execute'):
            raise ValueError('未知模式')
        plan = discover(requested)
        if mode == 'plan':
            print(json.dumps(plan, ensure_ascii=False))
        else:
            if len(sys.argv) != 5 or sys.argv[4] != 'ERASE_LVM_AND_WIPEFS':
                raise ValueError('缺少明确确认字符串')
            if not requested or plan['fingerprint'] != sys.argv[3]:
                raise ValueError('设备状态已变化或未选择目标;拒绝执行')
            print(json.dumps({'completed': execute(plan)}, ensure_ascii=False))
    except (ValueError, RuntimeError, OSError, subprocess.TimeoutExpired) as exc:
        print(str(exc), file=sys.stderr)
        sys.exit(1)

5. 验证与已知边界

生成笔记前,两个 Playbook 均已通过 ansible-playbook --syntax-check,两个 Python 文件均已通过 python3 -m py_compile。LVM role 的只读发现曾在用户提供的五台主机上成功运行;最新的 VirtIO /dev/vdX 支持和 always 清理逻辑尚未在这些主机上复测。健康巡检 role 尚未在目标 Linux 主机上实测。版本不同的 lsblk、LVM2 或设备映射可能使发现失败;失败时不会进入 LVM 删除步骤。

©著作权归作者所有,转载或内容合作请联系作者
【社区内容提示】社区部分内容疑似由AI辅助生成,浏览时请结合常识与多方信息审慎甄别。
平台声明:文章内容(如有图片或视频亦包括在内)由作者上传并发布,文章内容仅代表作者本人观点,简书系信息发布平台,仅提供信息存储服务。

友情链接更多精彩内容