基于Zabbix IPMI监控服务器硬件状况

简介:

公司有多个分部,且机房没有专业值班,机房等级不够。在这种情况下,又想实时监控机房环境,于是使用IPMI方式来达到目的。由于之前已经部署了Zabbix监控系统,本次将结合Zabbix自带的IPMI,完成服务器温度及风扇转速等的监控。

1.环境说明

被监控端服务器型号:Dell PowerEdge R510 
规划分配的IPMI地址: 10.103.1.100

2.Zabbix监控平台说明

Zabbix版本: 3.2.1,在安装时,未使用--with-openipmi 
Zabbix网络接口可以连通10.103.1.100

3.前置学习

维基百科IPMI: http://zh.wikipedia.org/wiki/IPMI 
IBM DeveloperWorks -- 使用ipmitool实现Linux系统下对服务器的ipmi管理:http://www.ibm.com/developerworks/cn/linux/l-ipmi/ 
Dell -- Managing Dell PowerEdge Servers Using IPMItool:http://www.dell.com/downloads/global/power/ps4q04-20040204-Murphy.pdf 
Zabbix IPMI checks:https://www.zabbix.com/documentation/3.2/manual/config/items/itemtypes/ipmi 
使用IPMITOOL实现终端重定向(课外读物):http://docs.linuxtone.org/ebooks/Dell/ipmitool.pdf

4.配置IPMI

4.1.配置IPMI地址

可以参考前置推荐中的《Managing Dell PowerEdge Servers Using IPMItool》在服务器启动时进行IPMI地址的配置,并开启IPMI Over LAN。 
也可以使用Dell的iDRAC开启IPMI功能,具体可以查看文章最后的参考资料。

2

4.2.获取传感器信息

登录Zabbix服务器,通过ipmitool远程访问Dell服务器传感器信息

# ipmitool -I lan -H 10.103.1.100 -U root -P calvin -L user sensor list
# ipmitool -I lan -H 10.103.1.100 -U root -P calvin -L user sensor get "FAN MOD 1B RPM"

2

2

4.3.安装IPMItool软件包

# yum -y install OpenIPMI OpenIPMI-devel ipmitool freeipmi

4.4.配置Zabbix

注:为了支持IPMI,需要在zabbix server/proxy安装时增加--with-openipmi参数

服务器端配置zabbix IPMI pollers 
zabbix_server.conf/zabbix_proxy.conf

# sed -i '/# StartIPMIPollers=0/aStartIPMIPollers=5' zabbix_server.conf
# service zabbix-server restart

4.5.导入监控模板

下面提供DELL的2个型号的IPMI模板: 
template-ipmi-dell-poweredge-r510 
template-ipmi-dell-poweredge-2950 
添加监控主机,关联上本模板,并在IPMI页面,设置Authentication algorithmDefault,Privilege levelUserUsernamesensorPasswordsensor_pass,保存即可。 
使用此种方法获取数据的结果就是效率很差,基本没什么数据。

5.使用Zabbix External checks自定义IPMI

本来是选择nagios的IPMI插件:check_ipmi_sensor,文件是:check_ipmi_sensor_v3-v3.9.tar.gz 
具体使用方法详见:http://www.thomas-krenn.com/en/wiki/IPMI_Sensor_Monitoring_Plugin

5.1.安装perl-IPC-Run模块

yum -y install perl-IPC-Run perl-Getopt-Long

5.2.使用check_ipmi_sensor查看效果

但是发现报错。

# ./check_ipmi_sensor -f ipmi.cfg -H 10.103.1.100 -vvv
------------- debug output for sel (-vvv is set): ------------
  /usr/sbin/ipmi-sel was executed with the following parameters:
    /usr/sbin/ipmi-sel -h 10.103.1.100 --config-file ipmi.cfg --driver-type=LAN_2_0 --output-event-state --interpret-oem-data --entity-sensor-names
  output of FreeIPMI:
ID  | Date        | Time     | Name                                        | Type                     | State    | Event
1   | Apr-08-2011 | 06:42:13 | System Board SEL                            | Event Logging Disabled   | Nominal  | Log Area Reset/Cleared
2   | Jan-01-1970 | 08:00:31 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
3   | Jan-01-1970 | 08:00:36 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
4   | Aug-15-2011 | 23:09:53 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
5   | Aug-16-2011 | 11:38:25 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
6   | Aug-16-2011 | 11:38:25 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
7   | Aug-16-2011 | 11:38:55 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
8   | Jun-10-2012 | 22:41:13 | System Board Ambient Temp                   | Temperature              | Warning  | Upper Non-critical - going high ; Sensor Reading = 45.00 C ; Threshold = 45.00 C
9   | Jun-11-2012 | 02:53:53 | System Board Ambient Temp                   | Temperature              | Nominal  | Upper Non-critical - going high ; Sensor Reading = 43.00 C ; Threshold = 45.00 C
10  | Nov-05-2012 | 21:56:42 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
11  | Nov-14-2012 | 21:53:58 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
12  | Nov-14-2012 | 21:53:58 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
13  | Nov-14-2012 | 21:54:19 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
14  | Nov-15-2012 | 16:12:03 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
15  | Nov-17-2012 | 17:14:34 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
16  | Nov-17-2012 | 17:14:34 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
17  | Nov-17-2012 | 17:15:40 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
18  | Nov-19-2012 | 20:47:57 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
19  | Nov-19-2012 | 20:50:04 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
20  | Jan-01-1970 | 08:00:33 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
21  | Jan-01-1970 | 08:00:38 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
22  | Jun-27-2014 | 17:27:38 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
23  | Jun-27-2014 | 17:27:53 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
24  | Jan-01-1970 | 08:00:31 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
25  | Jan-01-1970 | 08:00:36 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
26  | Oct-31-2016 | 05:48:35 | System Board Ambient Temp                   | Temperature              | Warning  | Lower Non-critical - going low ; Sensor Reading = 8.00 C ; Threshold = 8.00 C
27  | Oct-31-2016 | 09:00:38 | System Board Ambient Temp                   | Temperature              | Nominal  | Lower Non-critical - going low ; Sensor Reading = 10.00 C ; Threshold = 8.00 C
------------- debug output for sensors (-vvv is set): ------------
  script was executed with the following parameters:
    ./check_ipmi_sensor -f ipmi.cfg -H 10.103.1.100 -vvv
  check_ipmi_sensor version:
    3.9
  FreeIPMI version:
    ipmi-sensors - 1.2.9
  FreeIPMI was executed with the following parameters:
    /usr/sbin/ipmi-sensors -h 10.103.1.100 --config-file ipmi.cfg --quiet-cache --sdr-cache-recreate --interpret-oem-data --output-sensor-state --ignore-not-available-sensors --driver-type=LAN_2_0 --output-sensor-thresholds
  FreeIPMI return code: 0
  output of FreeIPMI:
Record ID | Sensor Name | Sensor Group | Monitoring Status | Sensor Units | Sensor Reading
5 | Ambient Temp | Temperature | Nominal | C | 28.000000
7 | CMOS Battery | Battery | Nominal | N/A | 'OK'
8 | VCORE PG | Voltage | Nominal | N/A | 'State Deasserted'
9 | VCORE PG | Voltage | Nominal | N/A | 'State Deasserted'
10 | 0.75 VTT PG | Voltage | Nominal | N/A | 'State Deasserted'
11 | 0.75 VTT PG | Voltage | Nominal | N/A | 'State Deasserted'
12 | CPU VTT PG | Voltage | Nominal | N/A | 'State Deasserted'
13 | 1.5V PG | Voltage | Nominal | N/A | 'State Deasserted'
14 | 1.8V PG | Voltage | Nominal | N/A | 'State Deasserted'
15 | 5V PG | Voltage | Nominal | N/A | 'State Deasserted'
16 | MEM CPU2 FAIL | Voltage | Nominal | N/A | 'State Deasserted'
17 | 5V Riser PG | Voltage | Nominal | N/A | 'State Deasserted'
18 | MEM CPU1 FAIL | Voltage | Nominal | N/A | 'State Deasserted'
19 | VTT CPU2 FAIL | Voltage | Nominal | N/A | 'State Deasserted'
20 | VTT CPU1 FAIL | Voltage | Nominal | N/A | 'State Deasserted'
21 | 0.9V PG | Voltage | Nominal | N/A | 'State Deasserted'
22 | CPU2 1.8 PLL PG | Voltage | Nominal | N/A | 'State Deasserted'
23 | CPU1 1.8 PLL PG | Voltage | Nominal | N/A | 'State Deasserted'
24 | 1.1 FAIL | Voltage | Nominal | N/A | 'State Deasserted'
25 | 1.0 LOM FAIL | Voltage | Nominal | N/A | 'State Deasserted'
26 | 1.0 AUX FAIL | Voltage | Nominal | N/A | 'State Deasserted'
27 | Heatsink Pres | Entity Presence | Nominal | N/A | 'Entity Present'
28 | iDRAC6 Ent Pres | Entity Presence | Critical | N/A | 'Entity Absent'
29 | USB Cable Pres | Entity Presence | Nominal | N/A | 'Entity Present'
31 | Riser Presence | Entity Presence | Nominal | N/A | 'Entity Present'
32 | FAN MOD 1A RPM | Fan | Nominal | RPM | 3480.000000
34 | FAN MOD 2A RPM | Fan | Nominal | RPM | 3480.000000
36 | FAN MOD 3A RPM | Fan | Nominal | RPM | 3480.000000
39 | FAN MOD 4A RPM | Fan | Nominal | RPM | 3480.000000
40 | Presence | Entity Presence | Nominal | N/A | 'Entity Present'
41 | Presence | Entity Presence | Nominal | N/A | 'Entity Present'
42 | Presence | Entity Presence | Nominal | N/A | 'Entity Present'
43 | Presence | Entity Presence | Nominal | N/A | 'Entity Present'
44 | Presence  | Entity Presence | Nominal | N/A | 'Entity Present'
45 | Status | Processor | Nominal | N/A | 'Processor Presence detected'
46 | Status | Processor | Nominal | N/A | 'Processor Presence detected'
47 | Status | Power Supply | Nominal | N/A | 'Presence detected'
48 | Current | Current | Nominal | A | 0.400000
49 | Current | Current | Nominal | A | 0.400000
50 | Voltage | Voltage | Nominal | V | 218.000000
51 | Voltage | Voltage | Nominal | V | 218.000000
52 | Status | Power Supply | Nominal | N/A | 'Presence detected'
53 | Status | Cable/Interconnect | Nominal | N/A | 'Cable/Interconnect is connected'
54 | OS Watchdog | Watchdog 2 | Nominal | N/A | 'OK'
56 | Intrusion | Physical Security | Nominal | N/A | 'OK'
57 | PS Redundancy | Power Supply | Nominal | N/A | 'Fully Redundant'
58 | Fan Redundancy | Fan | Nominal | N/A | 'Fully Redundant'
60 | System Level | Current | Nominal | W | 168.000000
61 | Power Optimized | OEM Reserved | Nominal | N/A | 'Good'
62 | Drive | Drive Slot | Nominal | N/A | 'Drive Presence'
65 | Cable SAS A | Cable/Interconnect | Nominal | N/A | 'Cable/Interconnect is connected'
66 | Cable SAS B | Cable/Interconnect | Nominal | N/A | 'Cable/Interconnect is connected'
67 | DKM Status | OEM Reserved | N/A | N/A | 'OEM Event = 0000h'
119 | FAN MOD 5A RPM | Fan | Nominal | RPM | 3480.000000

--------------------- end of debug output ---------------------
IPMI Status: Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 737.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 738.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 749.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in concatenation (.) or string at ./check_ipmi_sensor line 750.
Use of uninitialized value in string ne at ./check_ipmi_sensor line 759.
Critical [iDRAC6 Ent Pres = Critical ('Entity Absent'), System Board Intrusion = Critical (Physical Security), System Board Intrusion = Critical (Physical Security), Disk Drive Bay 1 Drive 2 = Critical (Drive Slot), Disk Drive Bay 1 Drive 2 = Critical (Drive Slot), System Board Ambient Temp = Warning (Temperature), Disk Drive Bay 1 Drive 2 = Critical (Drive Slot), Disk Drive Bay 1 Drive 2 = Critical (Drive Slot), Disk Drive Bay 1 Drive 2 = Critical (Drive Slot), Disk Drive Bay 1 Drive 2 = Critical (Drive Slot), System Board Intrusion = Critical (Physical Security), System Board Intrusion = Critical (Physical Security), System Board Intrusion = Critical (Physical Security), System Board Intrusion = Critical (Physical Security), System Board Ambient Temp = Warning (Temperature)] | 'Ambient Temp'=28.000000;:;: 'FAN MOD 1A RPM'=3480.000000;:;: 'FAN MOD 2A RPM'=3480.000000;:;: 'FAN MOD 3A RPM'=3480.000000;:;: 'FAN MOD 4A RPM'=3480.000000;:;: 'Current'=0.400000;:;: 'Current'=0.400000;:;: 'Voltage'=218.000000;:;: 'Voltage'=218.000000;:;: 'System Level'=168.000000;:;: 'FAN MOD 5A RPM'=3480.000000;:;:
Ambient Temp = 28.000000 (Status: Nominal)
CMOS Battery = 'OK' (Status: Nominal)
VCORE PG = 'State Deasserted' (Status: Nominal)
VCORE PG = 'State Deasserted' (Status: Nominal)
0.75 VTT PG = 'State Deasserted' (Status: Nominal)
0.75 VTT PG = 'State Deasserted' (Status: Nominal)
CPU VTT PG = 'State Deasserted' (Status: Nominal)
1.5V PG = 'State Deasserted' (Status: Nominal)
1.8V PG = 'State Deasserted' (Status: Nominal)
5V PG = 'State Deasserted' (Status: Nominal)
MEM CPU2 FAIL = 'State Deasserted' (Status: Nominal)
5V Riser PG = 'State Deasserted' (Status: Nominal)
MEM CPU1 FAIL = 'State Deasserted' (Status: Nominal)
VTT CPU2 FAIL = 'State Deasserted' (Status: Nominal)
VTT CPU1 FAIL = 'State Deasserted' (Status: Nominal)
0.9V PG = 'State Deasserted' (Status: Nominal)
CPU2 1.8 PLL PG = 'State Deasserted' (Status: Nominal)
CPU1 1.8 PLL PG = 'State Deasserted' (Status: Nominal)
1.1 FAIL = 'State Deasserted' (Status: Nominal)
1.0 LOM FAIL = 'State Deasserted' (Status: Nominal)
1.0 AUX FAIL = 'State Deasserted' (Status: Nominal)
Heatsink Pres = 'Entity Present' (Status: Nominal)
iDRAC6 Ent Pres = 'Entity Absent' (Status: Critical)
USB Cable Pres = 'Entity Present' (Status: Nominal)
Riser Presence = 'Entity Present' (Status: Nominal)
FAN MOD 1A RPM = 3480.000000 (Status: Nominal)
FAN MOD 2A RPM = 3480.000000 (Status: Nominal)
FAN MOD 3A RPM = 3480.000000 (Status: Nominal)
FAN MOD 4A RPM = 3480.000000 (Status: Nominal)
Presence = 'Entity Present' (Status: Nominal)
Presence = 'Entity Present' (Status: Nominal)
Presence = 'Entity Present' (Status: Nominal)
Presence = 'Entity Present' (Status: Nominal)
Presence = 'Entity Present' (Status: Nominal)
Status = 'Processor Presence detected' (Status: Nominal)
Status = 'Processor Presence detected' (Status: Nominal)
Status = 'Presence detected' (Status: Nominal)
Current = 0.400000 (Status: Nominal)
Current = 0.400000 (Status: Nominal)
Voltage = 218.000000 (Status: Nominal)
Voltage = 218.000000 (Status: Nominal)
Status = 'Presence detected' (Status: Nominal)
Status = 'Cable/Interconnect is connected' (Status: Nominal)
OS Watchdog = 'OK' (Status: Nominal)
Intrusion = 'OK' (Status: Nominal)
PS Redundancy = 'Fully Redundant' (Status: Nominal)
Fan Redundancy = 'Fully Redundant' (Status: Nominal)
System Level = 168.000000 (Status: Nominal)
Power Optimized = 'Good' (Status: Nominal)
Drive = 'Drive Presence' (Status: Nominal)
Cable SAS A = 'Cable/Interconnect is connected' (Status: Nominal)
Cable SAS B = 'Cable/Interconnect is connected' (Status: Nominal)
FAN MOD 5A RPM = 3480.000000 (Status: Nominal)不过根据它的提示(其实插件也是调用如下命令),可以使用

/usr/sbin/ipmi-sel -h 10.103.1.100 --config-file ipmi.cfg --driver-type=LAN_2_0 --output-event-state --interpret-oem-data --entity-sensor-names执行结果是:

# /usr/sbin/ipmi-sel -h 10.103.1.100 --config-file ipmi.cfg --driver-type=LAN_2_0 --output-event-state --interpret-oem-data --entity-sensor-names
ID  | Date        | Time     | Name                                        | Type                     | State    | Event
1   | Apr-08-2011 | 06:42:13 | System Board SEL                            | Event Logging Disabled   | Nominal  | Log Area Reset/Cleared
2   | Jan-01-1970 | 08:00:31 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
3   | Jan-01-1970 | 08:00:36 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
4   | Aug-15-2011 | 23:09:53 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
5   | Aug-16-2011 | 11:38:25 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
6   | Aug-16-2011 | 11:38:25 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
7   | Aug-16-2011 | 11:38:55 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
8   | Jun-10-2012 | 22:41:13 | System Board Ambient Temp                   | Temperature              | Warning  | Upper Non-critical - going high ; Sensor Reading = 45.00 C ; Threshold = 45.00 C
9   | Jun-11-2012 | 02:53:53 | System Board Ambient Temp                   | Temperature              | Nominal  | Upper Non-critical - going high ; Sensor Reading = 43.00 C ; Threshold = 45.00 C
10  | Nov-05-2012 | 21:56:42 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
11  | Nov-14-2012 | 21:53:58 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
12  | Nov-14-2012 | 21:53:58 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
13  | Nov-14-2012 | 21:54:19 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
14  | Nov-15-2012 | 16:12:03 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
15  | Nov-17-2012 | 17:14:34 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
16  | Nov-17-2012 | 17:14:34 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Critical | Drive Fault ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
17  | Nov-17-2012 | 17:15:40 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
18  | Nov-19-2012 | 20:47:57 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
19  | Nov-19-2012 | 20:50:04 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
20  | Jan-01-1970 | 08:00:33 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
21  | Jan-01-1970 | 08:00:38 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
22  | Jun-27-2014 | 17:27:38 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
23  | Jun-27-2014 | 17:27:53 | Disk Drive Bay 1 Drive 2                    | Drive Slot               | Nominal  | Drive Presence ; OEM Event Data2 code = 01h ; OEM Event Data3 code = 02h
24  | Jan-01-1970 | 08:00:31 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
25  | Jan-01-1970 | 08:00:36 | System Board Intrusion                      | Physical Security        | Critical | General Chassis Intrusion ; Intrusion while system Off
26  | Oct-31-2016 | 05:48:35 | System Board Ambient Temp                   | Temperature              | Warning  | Lower Non-critical - going low ; Sensor Reading = 8.00 C ; Threshold = 8.00 C
27  | Oct-31-2016 | 09:00:38 | System Board Ambient Temp                   | Temperature              | Nominal  | Lower Non-critical - going low ; Sensor Reading = 10.00 C ; Threshold = 8.00 C

5.3编写Zabbix外部检查(External checks)脚本

# pwd
/usr/local/zabbix/share/zabbix/externalscripts
# cat check_ipmi

下面是脚本内容

#!/bin/bash
#用于检测ipmi相关信息
#Create on 2016-011-18
#@author: Chinge_Yang

args="$*"
echo $(date +%F-%T) $args >> /tmp/check_ipmi.debug

check_ipmi_dir=/usr/local/zabbix/shell/check_ipmi_sensor
check_ipmi_bin=$check_ipmi_dir/check_ipmi_sensor

ipmi_sensors=/usr/sbin/ipmi-sensors
ipmi_cfg=$check_ipmi_dir/ipmi.cfg

#$check_ipmi_bin -f $ipmi_cfg -v $args
#${ipmi_sel} $args --config-file $ipmi_cfg --driver-type=LAN_2_0 --output-event-state --interpret-oem-data --entity-sensor-names 
options="--quiet-cache --sdr-cache-recreate --interpret-oem-data --output-sensor-state --ignore-not-available-sensors --driver-type=LAN_2_0 --output-sensor-thresholds"

function usage(){
    echo "Usage: `basename $0` options (-h HOST|-n NAME)"
}

function check(){
    result=$($ipmi_sensors -h $host --config-file $ipmi_cfg $options|grep "$name"|awk -F"| " '{print $NF}')
    printf "%.4f\n" $result
}

if [ $# -lt 4 ]  
then
    usage
    exit 55     
fi  

# 用法: scriptname -options
# 注意: 必须使用破折号 (-) 
# 参数后接冒号,表示必须接值
while getopts ":h:n:" Option;do
  case $Option in
    h)
    host=$OPTARG
    ;;
    n)
    name=$OPTARG
    ;;
    *)
    usage
    ;;   # 默认情况的处理
  esac
done

shift $(($OPTIND - 1))
#  (译者注: shift命令是可以带参数的, 参数就是移动的个数)
#  将参数指针减1, 这样它将指向下一个参数.
#  $1 现在引用的是命令行上的第一个非选项参数,
#+ 如果有一个这样的参数存在的话.

check

exit 0

添加执行权限

chmod a+x check_ipmi

5.4新建自定义模板

这里就不详细介绍内容了,其实就是改改上文中的模板而来,一张图看完内容:

2

给2张图看看效果:

2

2

好吧,最后发现,就算是自定义脚本,仍然是获取数据艰难,脚本执行ipmi的命令都timeout。。。。



本文转自 ygqygq2 51CTO博客,原文链接:http://blog.51cto.com/ygqygq2/1874277,如需转载请自行联系原作者

相关文章
|
17天前
|
监控 关系型数据库 MySQL
zabbix agent集成percona监控MySQL的插件实战案例
这篇文章是关于如何使用Percona监控插件集成Zabbix agent来监控MySQL的实战案例。
28 2
zabbix agent集成percona监控MySQL的插件实战案例
|
10天前
|
存储 弹性计算 运维
自动化监控和响应ECS系统事件
阿里云提供的ECS系统事件用于记录云资源信息,如实例启停、到期通知等。为实现自动化运维,如故障处理与动态调度,可使用云助手插件`ecs-tool-event`。该插件定时获取并转化ECS事件为日志存储,便于监控与响应,无需额外开发,适用于大规模集群管理。详情及示例可见链接文档。
|
16天前
|
存储 监控 Linux
监控Linux服务器
详细介绍了如何监控Linux服务器,包括监控CPU、内存、磁盘存储和带宽的使用情况,以及使用各种系统监控工具如vmstat、iostat、sar、top和dstat来分析系统性能,并推荐了一些开源监控系统。
24 0
监控Linux服务器
|
22天前
|
Prometheus 监控 Cloud Native
Web服务器的日志分析与监控
【8月更文第28天】Web服务器日志提供了关于服务器活动的重要信息,包括访问记录、错误报告以及性能数据。有效地分析这些日志可以帮助我们了解用户行为、诊断问题、优化网站性能,并确保服务的高可用性。本文将介绍如何使用日志分析和实时监控工具来监测Web服务器的状态和性能指标,并提供具体的代码示例。
114 0
|
27天前
|
监控 Linux 测试技术
|
1月前
|
机器学习/深度学习 编解码 人工智能
阿里云gpu云服务器租用价格:最新收费标准与活动价格及热门实例解析
随着人工智能、大数据和深度学习等领域的快速发展,GPU服务器的需求日益增长。阿里云的GPU服务器凭借强大的计算能力和灵活的资源配置,成为众多用户的首选。很多用户比较关心gpu云服务器的收费标准与活动价格情况,目前计算型gn6v实例云服务器一周价格为2138.27元/1周起,月付价格为3830.00元/1个月起;计算型gn7i实例云服务器一周价格为1793.30元/1周起,月付价格为3213.99元/1个月起;计算型 gn6i实例云服务器一周价格为942.11元/1周起,月付价格为1694.00元/1个月起。本文为大家整理汇总了gpu云服务器的最新收费标准与活动价格情况,以供参考。
阿里云gpu云服务器租用价格:最新收费标准与活动价格及热门实例解析
|
9天前
|
Cloud Native Java 编译器
将基于x86架构平台的应用迁移到阿里云倚天实例云服务器参考
随着云计算技术的不断发展,云服务商们不断推出高性能、高可用的云服务器实例,以满足企业日益增长的计算需求。阿里云推出的倚天实例,凭借其基于ARM架构的倚天710处理器,提供了卓越的计算能力和能效比,特别适用于云原生、高性能计算等场景。然而,有的用户需要将传统基于x86平台的应用迁移到倚天实例上,本文将介绍如何将基于x86架构平台的应用迁移到阿里云倚天实例的服务器上,帮助开发者和企业用户顺利完成迁移工作,享受更高效、更经济的云服务。
将基于x86架构平台的应用迁移到阿里云倚天实例云服务器参考
|
7天前
|
编解码 前端开发 安全
通过阿里云的活动购买云服务器时如何选择实例、带宽、云盘
在我们选购阿里云服务器的过程中,不管是新用户还是老用户通常都是通过阿里云的活动去买了,一是价格更加实惠,二是活动中的云服务器配置比较丰富,足可以满足大部分用户的需求,但是面对琳琅满目的云服务器实例、带宽和云盘选项,如何选择更适合自己,成为许多用户比较关注的问题。本文将介绍如何在阿里云的活动中选择合适的云服务器实例、带宽和云盘,以供参考和选择。
通过阿里云的活动购买云服务器时如何选择实例、带宽、云盘
|
6天前
|
弹性计算 运维 安全
阿里云轻量应用服务器和经济型e实例区别及选择参考
目前在阿里云的活动中,轻量应用服务器2核2G3M带宽价格为82元1年,2核2G3M带宽的经济型e实例云服务器价格99元1年,对于云服务器配置和性能要求不是很高的阿里云用户来说,这两款服务器配置和价格都差不多,阿里云轻量应用服务器和ECS云服务器让用户二选一,很多用户不清楚如何选择,本文来说说轻量应用服务器和经济型e实例的区别及选择参考。
阿里云轻量应用服务器和经济型e实例区别及选择参考
|
7天前
|
机器学习/深度学习 存储 人工智能
阿里云GPU云服务器实例规格gn6v、gn7i、gn6i实例性能及区别和选择参考
阿里云的GPU云服务器产品线在深度学习、科学计算、图形渲染等多个领域展现出强大的计算能力和广泛的应用价值。本文将详细介绍阿里云GPU云服务器中的gn6v、gn7i、gn6i三个实例规格族的性能特点、区别及选择参考,帮助用户根据自身需求选择合适的GPU云服务器实例。
阿里云GPU云服务器实例规格gn6v、gn7i、gn6i实例性能及区别和选择参考

推荐镜像

更多