Python 正则表达式实战之Java日志解析

本文涉及的产品
云数据库 RDS MySQL,集群系列 2核4GB
推荐场景:
搭建个人博客
RDS MySQL Serverless 基础系列,0.5-2RCU 50GB
日志服务 SLS,月写入数据量 50GB 1个月
简介: Python 正则表达式实战之Java日志解析

需求描述

基于生产监控告警需求,需要对Java日志进行解析,提取相关信息,作为告警通知消息的内容部分。

提取思路

具体怎么提取,提取哪些内容呢?这里笔者分析了大量不同形态的生产日志,最后总结出4种形态,如下,制定出以下提取逻辑。

形态1

上图中,款选部分即为要提取的主要内容,即异常发生时所在文件,代码行,自定义异常相关描述,异常类型,异常描述,这里提取的相关说明和异常描述将统一作为异常的详细描述

形态2

类似形态1,如果没有独占一行的“异常类型”,那就取最后Caused by:后面的异常类型,及其描述

形态3

形态1,形态2不匹配的情况下,匹配形态3,该形态中,异常类型和描述是包含在自定义异常相关描述里面的

形态4

前三者都不匹配的情况下,匹配最后这种形态。没有异常类型,仅日志级别“ERROR”可以标识它是条异常日志。

代码实现

#!/usr/bin/env python
#-*- coding:utf-8 -*-
import re
log_list = [
'''
2021-10-18 09:22:41,079:ERROR http-nio-9330-exec-4 (DirectJDKLog.java:181) - Servlet.service() for servlet [dispatcherServlet] in context with path [/finance] threw exception [Request processing failed; nested exception is java.lang.NullPointerException] with root cause
java.lang.NullPointerException
  at java.util.Comparator.lambda$comparing$77a9974f$1(Comparator.java:469) ~[?:1.8.0_202]
  at java.util.TreeMap.put(TreeMap.java:552) ~[?:1.8.0_202]
''',
'''
2021-10-18 09:22:55,222:WARN kafka-async-consumer-2 (FeignClientsErrorDecoder.java:43) - read Exception failed!
com.fasterxml.jackson.databind.JsonMappingException: No content to map due to end-of-input
    at [Source: java.io.InputStreamReader@743333a3; line: 1, column: 0]
  at com.fasterxml.jackson.databind.JsonMappingException.from(JsonMappingException.java:270) ~[jackson-databind-2.8.4.jar!/:2.8.4]
''',
'''
2021-10-18 09:22:52,975:ERROR [,] parallel-2 (AccessLogWebFilter.java:60) - [accessId=616ccc49ff642e00010a4e8c] 发生网关内部错误
org.springframework.web.server.ResponseStatusException: 504 GATEWAY_TIMEOUT "Response took longer than timeout: PT35S"; nested exception is org.springframework.cloud.gateway.support.TimeoutException: Response took longer than timeout: PT35S
  at org.springframework.cloud.gateway.filter.NettyRoutingFilter.lambda$filter$5(NettyRoutingFilter.java:211) ~[spring-cloud-gateway-core-2.1.3.RELEASE.jar!/:2.1.3.RELEASE]
  at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) [?:1.8.0_202]
  at java.lang.Thread.run(Thread.java:748) [?:1.8.0_202]
Caused by: org.springframework.cloud.gateway.support.TimeoutException: Response took longer than timeout: PT35S
    at
''',
'''
2021-10-18 09:22:41,905:WARN http-nio-8080-exec-60 (VehicleOeImpl.java:1000) - 批量更新第三方价格失败1---->
org.springframework.jdbc.BadSqlGrammarException:
### Error updating database.  Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException: Query was empty
### The error may involve com.cmall.ec.webapp.maindata.web.dao.vehicleOe.ThirdPartyOeMapper.updateList-Inline
### The error may involve com.cmall.ec.webapp.maindata.web.dao.vehicleOe.ThirdPartyOeMapper.updateList-Inline
### The error occurred while setting parameters
### SQL:
### Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException: Query was empty
; bad SQL grammar []; nested exception is com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException: Query was empty
  at org.springframework.jdbc.support.SQLExceptionSubclassTranslator.doTranslate(SQLExceptionSubclassTranslator.java:91)
  at org.springframework.jdbc.support.AbstractFallbackSQLExceptionTranslator.translate(AbstractFallbackSQLExceptionTranslator.java:73)
  at org.apache.tomcat.util.threads.TaskThread$WrappingRunnable.run(TaskThread.java:61)
  at java.lang.Thread.run(Thread.java:748)
Caused by: com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException: Query was empty
  at java.lang.reflect.Method.invoke(Method.java:498)
  at org.mybatis.spring.SqlSessionTemplate$SqlSessionInterceptor.invoke(SqlSessionTemplate.java:434)
  ... 70 more
''',
'''
2021-10-17 18:39:33,066:ERROR http-nio-10062-exec-34 TID: 962118fb93d345bc92af98499ad0f771.3235.16344671730621817 (DirectJDKLog.java:181) - Servlet.service() for servlet [dispatcherServlet] in context with path [/orders/seller] threw exception [Request processing failed; nested exception is com.cmall.commons.service.exception.HttpMessageException: [400]标名为空] with root cause
Exception: 标名为空
  at com.cmall.commons.utils.Assert.fail(Assert.java:553) ~[icec-cloud-commons-0.4.5.jar!/:?]
  at com.cmall.commons.utils.Assert.notBlank(Assert.java:112) ~[icec-cloud-commons-0.4.5.jar!/:?]
''',
'''
2021-10-18 09:22:23,849:ERROR http-nio-10030-exec-2 TID: ed41cdfb8d5d4953a713285802c56032.80.16345201436864709 (DicountAssembler.java:266) - 查询商品类优惠结果失败:DiscountProductRequest(companyId=IYl6MgdkiG9KoBMYUJo, userLoginId=5c385563ad996c47bf5f7ccd, provinceGeoId=CN-43, cityGeoId=284)
feign.FeignException: status 500 reading DiscountProductClient#listDiscountProducts(DiscountProductRequest); content:
{"timestamp":1634520143844,"status":500,"error":"Internal Server Error","exception":"com.cmall.commons.service.exception.HttpMessageException","message":"优惠前的不含税价格不能为空","path":"/discountPromotion/listDiscountProducts"}
  at feign.FeignException.errorStatus(FeignException.java:62) ~[feign-core-9.3.1.jar!/:?]
  at java.lang.Thread.run(Thread.java:748) [?:1.8.0_202]
''',
'''2021-10-18 09:22:23,849:ERROR [-]  (DicountAssembler.java:266) - task supervisor threw an exception ....''',
'''
2021-10-18 09:13:13,940:ERROR kafka-async-consumer-9 (ConsumeSupport.java:104) - kafka消费失败, cid:616cca245cc0a90001ead690, message:{"jsonMessageType":"com.cmall.ec.cloud.scheduletask.values.kafka.command.DistributedDelayCommand","id":"616cca24d52d1d00010531dd","usage":"EVENT","service":"schedule-task-service","topic":"prod-quote-command-delay","timeStamp":1634519588985,"delayTimes":16983,"producerTaskId":"B21101807668","inquiryId":"B21101807668","resolveBatchId":"616cca019cf7b70001c35477","resolveIds":["616cca00d52d1d000138be9a"],"type":"AUTO","retryCount":2,"producer":"DistributedDelayCommand"}, cause:java.lang.reflect.InvocationTargetException
  at sun.reflect.GeneratedMethodAccessor4946.invoke(Unknown Source)
  at java.lang.reflect.Method.invoke(Method.java:498)
  at java.lang.Thread.run(Thread.java:748)
Caused by: com.cmall.messagebus.exception.MessageBaseException: 系统自动报价失败-->org.springframework.dao.DuplicateKeyException:
### Error updating database.  Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException: Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'
### The error may involve defaultParameterMap
### The error occurred while setting parameters
### Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException: Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'
; SQL []; Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'; nested exception is com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException: Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'
  at com.cmall.ec.cloud.quotation.handler.QuotationAllocationHandler.intelligentAutoQuoteDispatcher(QuotationAllocationHandler.java:80)
  ... 9 more
''',
'''
2021-10-16 14:37:19,951:ERROR DiscoveryClient-1 (TimedSupervisorTask.java:79) - task supervisor threw an exception
java.lang.OutOfMemoryError: Java heap space
''',
'''2021-10-16 14:37:19,951:ERROR DiscoveryClient-1 (TimedSupervisorTask.java:79) - task supervisor threw an exception
java.lang.OutOfMemoryError: Java heap space
at''',
'''
2021-10-23 14:03:00,785:ERROR kafka-async-consumer-1 (ConsumeSupport.java:104) - kafka消费失败, cid:6173a593f90d4200010ce3fe, message:{"jsonMessageType":"com.cmall.ec.cloud.events.order.OrderSent","id":"6173a593cb769e0001402a3b","usage":"EVENT","service":"orders-service","topic":"prod-order","timeStamp":1634968979938,"orderId":"S2110230001835","shipmentType":"LOGISTICS","logisticsCompany":"快送","shipmentNum":"","shipGroupId":"6173a59355c35d00014c390c","remark":"","productStoreId":"SZQD0001","userLoginId":"5e0aa6649cf5260001336437","username":"许庆杰","items":[{"productId":"00001","quantity":1}],"isLogisticsFeePayOnLine":true}, cause:java.lang.reflect.InvocationTargetException
  at sun.reflect.GeneratedMethodAccessor1909.invoke(Unknown Source)
  at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
  at java.lang.reflect.Method.invoke(Method.java:498)
  at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
  at java.lang.Thread.run(Thread.java:748)
Caused by: com.baomidou.mybatisplus.exceptions.MybatisPlusException: Error: Cannot execute insertBatch Method. Cause
  at com.baomidou.mybatisplus.service.impl.ServiceImpl.insertBatch(ServiceImpl.java:137)
  at org.springframework.aop.framework.ReflectiveMethodInvocation.proceed(ReflectiveMethodInvocation.java:179)
  ... 9 more
Caused by: org.apache.ibatis.exceptions.PersistenceException:
### Error flushing statements.  Cause: org.apache.ibatis.executor.BatchExecutorException: com.cmall.ec.cloud.service.dao.mapper.ShipFeeMapper.insert (batch index #1) failed. Cause: java.sql.BatchUpdateException: Duplicate entry 'SF2110230005145' for key 'PRIMARY'
### Cause: org.apache.ibatis.executor.BatchExecutorException: com.cmall.ec.cloud.service.dao.mapper.ShipFeeMapper.insert (batch index #1) failed. Cause: java.sql.BatchUpdateException: Duplicate entry 'SF2110230005145' for key 'PRIMARY'
  at org.apache.ibatis.exceptions.ExceptionFactory.wrapException(ExceptionFactory.java:30)
  ... 51 more
Caused by: org.apache.ibatis.executor.BatchExecutorException: com.cmall.ec.cloud.service.dao.mapper.ShipFeeMapper.insert (batch index #1) failed. Cause: java.sql.BatchUpdateException: Duplicate entry 'SF2110230005145' for key 'PRIMARY'
  at java.lang.reflect.Method.invoke(Method.java:498)
  ... 52 more
Caused by: java.sql.BatchUpdateException: Duplicate entry 'SF2110230005145' for key 'PRIMARY'
  at sun.reflect.GeneratedConstructorAccessor1540.newInstance(Unknown Source)
  ... 61 more
Caused by: com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException: Duplicate entry 'SF2110230005145' for key 'PRIMARY'
  at sun.reflect.GeneratedConstructorAccessor1537.newInstance(Unknown Source)
  ... 71 more
''',
'''
2021-10-25 16:29:15,853:ERROR reactor-http-epoll-3 (CompositeLog.java:122) - 500 Server Error for HTTP POST "/job-service/api/registry"
reactor.netty.http.client.PrematureCloseException: Connection prematurely closed BEFORE response
''',
'''2021-11-06 07:09:04,781:WARN http-nio-10011-exec-85 (AdminService.java:97) - 任务执行失败, code:500, reason:java.lang.OutOfMemoryError: Java heap space''',
'''2022-01-08 13:46:30,668:ERROR http-nio-9524-exec-5 (WaitSettleDealServiceImpl.java:147) - 添加退订单到账单失败com.cmall.commons.service.exception.HttpMessageException: [411]退货单添加失败,该分组已经结束对账,无法添加退货单'''
]
exception_match_pattern_list = [
    ':(ERROR|WARN) .+\s\(([^\s]+?\.java):(\d+)\)(.*)\n([^:\s>\u4e00-\u9fa5]*Exception|[^:\s>\u4e00-\u9fa5]*Error)(.*?)(\s+at\s|$)',
    ':(ERROR|WARN) .+\s\(([^\s]+?\.java):(\d+)\).*([^\n]*Caused by: )([^:\s>]*?Exception|[^:\s]*?Error)(.*?)(\s+at\s*|$)',
    ':(ERROR|WARN) .+\s\(([^\s]+?\.java):(\d+)\)([^\n]*?)([^:\s>\u4e00-\u9fa5]*Exception|[^:\s>\u4e00-\u9fa5]*Error)([^\n]*?)(\s+at\s|$)',
    ':(ERROR|WARN) .+\s\(([^\s]+?\.java):(\d+)\)(.*?)\n*?\s*?([^:\s>\u4e00-\u9fa5]*Exception|[^:\s>\u4e00-\u9fa5]*Error)*?(.*?)(\s+at\s|$)'
]
for log_index, log in enumerate(log_list):
    flag = 0
    for pattern_index, flag_pattern in enumerate(exception_match_pattern_list):
        match_result = re.findall(flag_pattern, log, re.DOTALL)
        if match_result:
            print('匹配第%s个Pattern' % (pattern_index+1), '匹配结果:', match_result[0])
            flag = 1
            break
    if not flag:
        print('第%s条日志,不匹配任何正则表达式' % (log_index + 1))

提取效果

匹配第1个Pattern 匹配结果: ('ERROR', 'DirectJDKLog.java', '181', ' - Servlet.service() for servlet [dispatcherServlet] in context with path [/finance] threw exception [Request processing failed; nested exception is java.lang.NullPointerException] with root cause', 'java.lang.NullPointerException', '', '\n\tat ')
匹配第1个Pattern 匹配结果: ('WARN', 'FeignClientsErrorDecoder.java', '43', ' - read Exception failed!', 'com.fasterxml.jackson.databind.JsonMappingException', ': No content to map due to end-of-input', '\n    at ')
匹配第1个Pattern 匹配结果: ('ERROR', 'AccessLogWebFilter.java', '60', ' - [accessId=616ccc49ff642e00010a4e8c] 发生网关内部错误', 'org.springframework.web.server.ResponseStatusException', ': 504 GATEWAY_TIMEOUT "Response took longer than timeout: PT35S"; nested exception is org.springframework.cloud.gateway.support.TimeoutException: Response took longer than timeout: PT35S', '\n\tat ')
匹配第1个Pattern 匹配结果: ('WARN', 'VehicleOeImpl.java', '1000', ' - 批量更新第三方价格失败1---->', 'org.springframework.jdbc.BadSqlGrammarException', ':\n### Error updating database.  Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException: Query was empty\n### The error may involve com.cmall.ec.webapp.maindata.web.dao.vehicleOe.ThirdPartyOeMapper.updateList-Inline\n### The error may involve com.cmall.ec.webapp.maindata.web.dao.vehicleOe.ThirdPartyOeMapper.updateList-Inline\n### The error occurred while setting parameters\n### SQL:\n### Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException: Query was empty\n; bad SQL grammar []; nested exception is com.mysql.jdbc.exceptions.jdbc4.MySQLSyntaxErrorException: Query was empty', '\n\tat ')
匹配第1个Pattern 匹配结果: ('ERROR', 'DirectJDKLog.java', '181', ' - Servlet.service() for servlet [dispatcherServlet] in context with path [/orders/seller] threw exception [Request processing failed; nested exception is com.cmall.commons.service.exception.HttpMessageException: [400]标名为空] with root cause', 'Exception', ': 标名为空', '\n\tat ')
匹配第1个Pattern 匹配结果: ('ERROR', 'DicountAssembler.java', '266', ' - 查询商品类优惠结果失败:DiscountProductRequest(companyId=IYl6MgdkiG9KoBMYUJo, userLoginId=5c385563ad996c47bf5f7ccd, provinceGeoId=CN-43, cityGeoId=284)', 'feign.FeignException', ': status 500 reading DiscountProductClient#listDiscountProducts(DiscountProductRequest); content:\n{"timestamp":1634520143844,"status":500,"error":"Internal Server Error","exception":"com.cmall.commons.service.exception.HttpMessageException","message":"优惠前的不含税价格不能为空","path":"/discountPromotion/listDiscountProducts"}', '\n\tat ')
匹配第4个Pattern 匹配结果: ('ERROR', 'DicountAssembler.java', '266', '', '', ' - task supervisor threw an exception ....', '')
匹配第2个Pattern 匹配结果: ('ERROR', 'ConsumeSupport.java', '104', 'Caused by: ', 'com.cmall.messagebus.exception.MessageBaseException', ": 系统自动报价失败-->org.springframework.dao.DuplicateKeyException:\n### Error updating database.  Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException: Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'\n### The error may involve defaultParameterMap\n### The error occurred while setting parameters\n### Cause: com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException: Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'\n; SQL []; Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'; nested exception is com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException: Duplicate entry '616cca225cc0a90001e1d0d2' for key 'PRIMARY'", '\n\tat ')
匹配第1个Pattern 匹配结果: ('ERROR', 'TimedSupervisorTask.java', '79', ' - task supervisor threw an exception', 'java.lang.OutOfMemoryError', ': Java heap space', '')
匹配第1个Pattern 匹配结果: ('ERROR', 'TimedSupervisorTask.java', '79', ' - task supervisor threw an exception', 'java.lang.OutOfMemoryError', ': Java heap space\nat', '')
匹配第2个Pattern 匹配结果: ('ERROR', 'ConsumeSupport.java', '104', 'Caused by: ', 'com.mysql.jdbc.exceptions.jdbc4.MySQLIntegrityConstraintViolationException', ": Duplicate entry 'SF2110230005145' for key 'PRIMARY'", '\n\tat ')
匹配第1个Pattern 匹配结果: ('ERROR', 'CompositeLog.java', '122', ' - 500 Server Error for HTTP POST "/job-service/api/registry"', 'reactor.netty.http.client.PrematureCloseException', ': Connection prematurely closed BEFORE response', '')
匹配第3个Pattern 匹配结果: ('WARN', 'AdminService.java', '97', ' - 任务执行失败, code:500, reason:', 'java.lang.OutOfMemoryError', ': Java heap space', '')
匹配第3个Pattern 匹配结果: ('ERROR', 'WaitSettleDealServiceImpl.java', '147', ' - 添加退订单到账单失败', 'com.cmall.commons.service.exception.HttpMessageException', ': [411]退货单添加失败,该分组已经结束对账,无法添加退货单', '')
相关实践学习
日志服务之使用Nginx模式采集日志
本文介绍如何通过日志服务控制台创建Nginx模式的Logtail配置快速采集Nginx日志并进行多维度分析。
目录
相关文章
|
15天前
|
XML JSON Java
Java中Log级别和解析
日志级别定义了日志信息的重要程度,从低到高依次为:TRACE(详细调试)、DEBUG(开发调试)、INFO(一般信息)、WARN(潜在问题)、ERROR(错误信息)和FATAL(严重错误)。开发人员可根据需要设置不同的日志级别,以控制日志输出量,避免影响性能或干扰问题排查。日志框架如Log4j 2由Logger、Appender和Layout组成,通过配置文件指定日志级别、输出目标和格式。
|
1月前
|
运维 Shell 数据库
Python执行Shell命令并获取结果:深入解析与实战
通过以上内容,开发者可以在实际项目中灵活应用Python执行Shell命令,实现各种自动化任务,提高开发和运维效率。
65 20
|
1月前
|
存储 Java 计算机视觉
Java二维数组的使用技巧与实例解析
本文详细介绍了Java中二维数组的使用方法
52 15
|
6天前
|
Java API 数据处理
深潜数据海洋:Java文件读写全面解析与实战指南
通过本文的详细解析与实战示例,您可以系统地掌握Java中各种文件读写操作,从基本的读写到高效的NIO操作,再到文件复制、移动和删除。希望这些内容能够帮助您在实际项目中处理文件数据,提高开发效率和代码质量。
15 0
|
2月前
|
人工智能 自然语言处理 Java
FastExcel:开源的 JAVA 解析 Excel 工具,集成 AI 通过自然语言处理 Excel 文件,完全兼容 EasyExcel
FastExcel 是一款基于 Java 的高性能 Excel 处理工具,专注于优化大规模数据处理,提供简洁易用的 API 和流式操作能力,支持从 EasyExcel 无缝迁移。
281 9
FastExcel:开源的 JAVA 解析 Excel 工具,集成 AI 通过自然语言处理 Excel 文件,完全兼容 EasyExcel
|
1月前
|
供应链 搜索推荐 API
深度解析1688 API对电商的影响与实战应用
在全球电子商务迅猛发展的背景下,1688作为知名的B2B电商平台,为中小企业提供商品批发、分销、供应链管理等一站式服务,并通过开放的API接口,为开发者和电商企业提供数据资源和功能支持。本文将深入解析1688 API的功能(如商品搜索、详情、订单管理等)、应用场景(如商品展示、搜索优化、交易管理和用户行为分析)、收益分析(如流量增长、销售提升、库存优化和成本降低)及实际案例,帮助电商从业者提升运营效率和商业收益。
214 20
|
1月前
|
算法 搜索推荐 Java
【潜意识Java】深度解析黑马项目《苍穹外卖》与蓝桥杯算法的结合问题
本文探讨了如何将算法学习与实际项目相结合,以提升编程竞赛中的解题能力。通过《苍穹外卖》项目,介绍了订单配送路径规划(基于动态规划解决旅行商问题)和商品推荐系统(基于贪心算法)。这些实例不仅展示了算法在实际业务中的应用,还帮助读者更好地准备蓝桥杯等编程竞赛。结合具体代码实现和解析,文章详细说明了如何运用算法优化项目功能,提高解决问题的能力。
68 6
|
1月前
|
SQL Java 数据库连接
如何在 Java 代码中使用 JSqlParser 解析复杂的 SQL 语句?
大家好,我是 V 哥。JSqlParser 是一个用于解析 SQL 语句的 Java 库,可将 SQL 解析为 Java 对象树,支持多种 SQL 类型(如 `SELECT`、`INSERT` 等)。它适用于 SQL 分析、修改、生成和验证等场景。通过 Maven 或 Gradle 安装后,可以方便地在 Java 代码中使用。
326 11
|
1月前
|
存储 算法 搜索推荐
【潜意识Java】期末考试可能考的高质量大题及答案解析
Java 期末考试大题整理:设计一个学生信息管理系统,涵盖面向对象编程、集合类、文件操作、异常处理和多线程等知识点。系统功能包括添加、查询、删除、显示所有学生信息、按成绩排序及文件存储。通过本题,考生可以巩固 Java 基础知识并掌握综合应用技能。代码解析详细,适合复习备考。
25 4
|
1月前
|
存储 分布式计算 Hadoop
基于Java的Hadoop文件处理系统:高效分布式数据解析与存储
本文介绍了如何借鉴Hadoop的设计思想,使用Java实现其核心功能MapReduce,解决海量数据处理问题。通过类比图书馆管理系统,详细解释了Hadoop的两大组件:HDFS(分布式文件系统)和MapReduce(分布式计算模型)。具体实现了单词统计任务,并扩展支持CSV和JSON格式的数据解析。为了提升性能,引入了Combiner减少中间数据传输,以及自定义Partitioner解决数据倾斜问题。最后总结了Hadoop在大数据处理中的重要性,鼓励Java开发者学习Hadoop以拓展技术边界。
61 7