用R读取PDF并进行数据挖掘

简介: 用R读取PDF并进行数据挖掘,例子如下: [javascript] view plaincopyprint? # here is a pdf for mining   url "http://www.

用R读取PDF并进行数据挖掘,例子如下:


[javascript]   view plain copy print ?
  1. # here is a pdf for mining  
  2. url "http://www.noisyroom.net/blog/RomneySpeech072912.pdf"  
  3. dest ".pdf")  
  4. download.file(url, dest, mode = "wb")  
  5.  
  6. # set path to pdftotxt.exe and convert pdf to text  
  7. exe "C:\\Program Files\\xpdfbin-win-3.03\\bin32\\pdftotext.exe"  
  8. system(paste("\"", exe, "\" \"", dest, "\"", sep = ""), wait = F)  
  9.  
  10. # get txt-file name and open it  
  11. filetxt ".pdf", ".txt", dest)  
  12. shell.exec(filetxt); shell.exec(filetxt) # strangely the first try always throws an error..  
  13.  
  14. # do something with it, i.e. a simple word cloud  
  15. library(tm)  
  16. library(wordcloud)  
  17. library(Rstem)  
  18.   
  19. txt 
  20.   
  21. txt 
  22. txt "\\f", stopwords()))  
  23.   
  24. corpus 
  25. corpus 
  26. tdm 
  27. m 
  28. d 
  29.  
  30. # Stem words  
  31. d$stem "english")  
  32.  
  33. # and put words to column, otherwise they would be lost when aggregating  
  34. d$word 
  35.  
  36. # remove web address (very long string):  
  37. d 
  38.  
  39. # aggregate freqeuncy by word stem and  
  40. # keep first words..  
  41. agg_freq 
  42. agg_word function(x) x[1])  
  43.   
  44. d 
  45.  
  46. # sort by frequency  
  47. d 
  48.  
  49. # print wordcloud:  
  50. wordcloud(d$word, d$freq)  
  51.  
  52. # remove files  
  53. file.remove(dir(tempdir(), full.name=T)) # remove files  

目录
相关文章
|
编解码 安全 Unix
数据导入与预处理-第4章-数据获取python读取pdf文档
数据导入与预处理-第4章-数据获取Python读取PDF文档 1 PDF简介 1.1 pdf是什么 2 Python操作PDF 2.1 pdfplumber库
数据导入与预处理-第4章-数据获取python读取pdf文档
|
自然语言处理 Python
Python 操作pdf文件(pdfplumber读取PDF写入Excel)
学习了解Python 操作pdf文件(pdfplumber读取PDF写入Excel)。
1401 0
Python 操作pdf文件(pdfplumber读取PDF写入Excel)
UIWebView 读取pdf,word,excel
UIWebView 读取pdf,word,excel
252 0
|
存储 Linux Python
Python编程:读取pdf、pptx、docx、xlsx文件的页数
Python编程:读取pdf、pptx、docx、xlsx文件的页数
1334 0
|
Python
python通过pdfminer或pdfminer3k读取pdf文件
python通过pdfminer或pdfminer3k读取pdf文件
533 0
|
Python
python通过pdfminer或pdfminer3k读取pdf文件
python通过pdfminer或pdfminer3k读取pdf文件
712 0
|
API .NET 开发框架
【Win10 开发】读取PDF文档
原文:【Win10 开发】读取PDF文档 关于用来读取PDF文档的内容的API,其实在Win8.1的时候就有,不过没关系,既咱们讨论的是10的UAP,连同8.1的内容也包括进去,所以老周无数次强调:把以前的内容学好了,就可以在不学习任何新知识的前提直接进入10的开发,至于你信不信,反正我信了。
1137 0