开发者社区> 问答> 正文

flink1.8.0 on yarn 程序运行一段时间报错,有什么解决方式吗?

flink1.8.0 on yarn 程序运行一段时间报如下错误,导致 The heartbeat of TaskManager with id container_1572430463280_50994_01_000004 timed out. 最终程序重启。

各位有没有碰到类似的问题,有什么解决方式吗?

jobmanager.log

2020-08-17 02:53:21,593 ERROR akka.remote.Remoting - Association to [akka.tcp://flink@${HOSTNAME}:36968] with UID [19 99537927] irrecoverably failed. Quarantining address. java.util.concurrent.TimeoutException: Remote system has been silent for too long. (more than 48.0 hours) at akka.remote.ReliableDeliverySupervisor$$anonfun$idle$1.applyOrElse(Endpoint.scala:375) at akka.actor.Actor$class.aroundReceive(Actor.scala:502) at akka.remote.ReliableDeliverySupervisor.aroundReceive(Endpoint.scala:203) at akka.actor.ActorCell.receiveMessage(ActorCell.scala:526) at akka.actor.ActorCell.invoke(ActorCell.scala:495) at akka.dispatch.Mailbox.processMailbox(Mailbox.scala:257) at akka.dispatch.Mailbox.run(Mailbox.scala:224) at akka.dispatch.Mailbox.exec(Mailbox.scala:234) at scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260) at scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339) at scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979) at scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)

*来自志愿者整理的flink邮件归档

展开
收起
游客sadna6pkvqnz6 2021-12-07 17:31:04 1346 0
1 条回答
写回答
取消 提交回答
  • container timeout 可以看下是不是 GC 的原因,看一下超时的这个 container container_1572430463280_50994_01_000004 的之前的 GC 情况*来自志愿者整理的flink

    2021-12-07 21:11:54
    赞同 展开评论 打赏
问答排行榜
最热
最新

相关电子书

更多
深度学习+大数据 TensorFlow on Yarn 立即下载
Docker on Yarn 微服务实践 立即下载
深度学习+大数据-TensorFlow on Yarn 立即下载