task blocked for more than 120 seconds

简介: ul  3 20:41:24 yz384 kernel:Jul  3 20:43:24 yz384 kernel: INFO: task chown:18647 blocked for more than 120 seconds.



ul  3 20:41:24 yz384 kernel:
Jul  3 20:43:24 yz384 kernel: INFO: task chown:18647 blocked for more than 120 seconds.
Jul  3 20:43:24 yz384 kernel: "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
Jul  3 20:43:24 yz384 kernel: chown         D ffffffff80157c4c     0 18647      1         24429 22803 (NOTLB)
Jul  3 20:43:24 yz384 kernel:  ffff81015ed4fde8 0000000000000086 ffff810c3fc1d3c0 ffff810e91a379c0
Jul  3 20:43:24 yz384 kernel:  0000000000000000 0000000000000008 ffff8103bf82f7a0 ffff810c3fd2e7a0
Jul  3 20:43:24 yz384 kernel:  000a29c122906514 000000000000141d ffff8103bf82f988 0000000a00000001
Jul  3 20:43:24 yz384 kernel: Call Trace:
Jul  3 20:43:24 yz384 kernel:  [<ffffffff80063c63>] __mutex_lock_slowpath+0x60/0x9b
Jul  3 20:43:24 yz384 kernel:  [<ffffffff80063cad>] .text.lock.mutex+0xf/0x14
Jul  3 20:43:24 yz384 kernel:  [<ffffffff8003b7e5>] chown_common+0x90/0xb0
Jul  3 20:43:24 yz384 kernel:  [<ffffffff80023fa0>] __user_walk_fd+0x41/0x4c
Jul  3 20:43:24 yz384 kernel:  [<ffffffff800e3c44>] sys_lchown+0x38/0x53
Jul  3 20:43:24 yz384 kernel:  [<ffffffff8000d57f>] dput+0x2c/0x114
Jul  3 20:43:24 yz384 kernel:  [<ffffffff80012d27>] __fput+0x191/0x1bd
Jul  3 20:43:24 yz384 kernel:  [<ffffffff8002d0f1>] mntput_no_expire+0x19/0x89
Jul  3 20:43:24 yz384 kernel:  [<ffffffff80024200>] filp_close+0x5c/0x64
Jul  3 20:43:24 yz384 kernel:  [<ffffffff8005d116>] system_call+0x7e/0x83
Jul  3 20:43:24 yz384 kernel:




The warning is given to indicate a problem with the system. In my experience it means that the process is blocked in kernel space for at least 120 seconds usually because the process is starved of disk I/O. This can be because of heavy swapping due to too much memory being used, e.g. if you have a heavy webserver load and you've configured too many apache child processes for your system. In your case it may just be that there are too many mysql processes competing for memory and data IO.


It can also happen if the underlying storage system is not performing well, e.g. if you have a SAN which is overloaded, or if there are soft errors on a disk which cause a lot of retries. Whenever a task has to wait long for its IO commands to complete, these warning may be issued.



This is a know bug. By default Linux uses up to 40% of the available memory for file system caching. After this mark has been reached the file system flushes all outstanding data to disk causing all following IOs going synchronous. For flushing out this data to disk this there is a time limit of 120 seconds by default. In the case here the IO subsystem is not fast enough to flush the data withing 120 seconds. This especially happens on systems with a lof of memory.


The problem is solved in later kernels and there is not “fix” from Oracle. I fixed this by lowering the mark for flushing the cache from 40% to 10% by setting “vm.dirty_ratio=10” in /etc/sysctl.conf. This setting does not influence overall database performance since you hopefully use Direct IO and bypass the file system cache completely.


原理:linux会设置40%的可用内存用来做系统cache,当flush数据时这40%内存中的数据由于和IO同步问题导致超时(120s),所将40%减小到10%,避免超时

在文件/etc/sysctl.conf中加入 vm.dirty_ratio=10



进程等待IO时,经常处于D状态,即TASK_UNINTERRUPTIBLE状态,处于这种状态的进程不处理信号,所以kill不掉,如果进程长期处于D状态,那么肯定不正常,原因可能有二:

1)IO路径上的硬件出问题了,比如硬盘坏了(只有少数情况会导致长期D,通常会返回错误);

2)内核自己出问题了。

这种问题一旦出现就通常不可恢复,kill不掉,通常只能重启恢复了。

内核针对这种开发了一种hung task的检测机制,基本原理是:定时检测系统中处于D状态的进程,如果其处于D状态的时间超过了指定时间(默认120s,可以配置),则打印相关堆栈信息,也可以通过proc参数配置使其直接panic。


目录
相关文章
|
移动开发 前端开发 CDN
移动端H5引入vconsole进行调试
移动端H5引入vconsole进行调试
1937 0
|
传感器 安全 API
SCP Firmware入门一篇就够啦
SCP Firmware入门一篇就够啦
2086 0
|
6月前
|
人工智能 自然语言处理 安全
【详细版教程】OpenClaw 2.6.2 Windows10 部署完整教程
OpenClaw(小龙虾)是专为Windows10优化的本地AI智能体框架,支持一键部署、零代码配置,可跨软件自动执行文件整理、邮件发送、浏览器操作等任务。数据完全本地运行,保障隐私安全,10分钟即可搭建属于你的AI数字员工。
|
10月前
|
Java Maven 数据安全/隐私保护
Nexus仓库
Nexus仓库是Sonatype推出的开源制品管理工具,支持Maven、Npm、Docker等格式。本文介绍其在Linux和Docker环境下的安装配置,包括JDK部署、OSS版下载、用户权限、匿名访问设置,以及仓库创建与上传下载操作,涵盖密码重置、数据持久化及脚本批量导入等内容,助力搭建高效私有仓库。
|
存储 Linux 编译器
Linux 交叉编译第三方库需要设置的环境变量
Linux 交叉编译第三方库需要设置的环境变量
1244 0
|
关系型数据库 应用服务中间件 nginx
基于Docker的LNMP环境微服务搭建
基于Docker的LNMP环境微服务搭建
基于Docker的LNMP环境微服务搭建
|
JavaScript 前端开发
使用 WebGL 创建 3D 动画
【10月更文挑战第3天】使用 WebGL 创建 3D 动画
|
机器学习/深度学习 缓存 自然语言处理
一文揭秘|预训练一个72b模型需要多久?
本文讲述评估和量化训练大规模语言模型,尤其是Qwen2-72B模型,所需的时间、资源和计算能力。
2301 12
|
缓存 安全 Unix
Linux 内核黑客不可靠指南【ChatGPT】
Linux 内核黑客不可靠指南【ChatGPT】
|
开发工具
ubuntuserver 修改networkd 为networkmanager
ubuntuserver 修改networkd 为networkmanager
847 0

热门文章

最新文章