Jupyter for Data Science
图书信息
| 作者 | Dan Toomey |
| 出版社 | Packt Publishing |
| ISBN | 9781785883293 |
| 出版时间 | 2017-10-20 |
| 字数 | 19.5万 |
| 分类 | 进口书,外文原版书,电脑,网络 |
读书简介
Your one-stop guide to building an efficient data science pipeline using Jupyter About This Book ? Get the most out of your Jupyter notebook to complete the trickiest of tasks in Data Science ? Learn all the tasks in the data science pipeline—from data acquisition to visualization—and implement them using Jupyter ? Get ahead of the curve by mastering all the applications of Jupyter for data science with this unique and intuitive guide Who This Book Is For This book targets students and professionals who wish to master the use of Jupyter to perform a variety of data science tasks. Some programming experience with R or Python, and some basic understanding of Jupyter, is all you need to get started with this book. What You Will Learn ? Understand why Jupyter notebooks are a perfect fit for your data science tasks ? Perform scientific computing and data analysis tasks with Jupyter ? Interpret and explore different kinds of data visually with charts, histograms, and more ? Extend SQL's capabilities with Jupyter notebooks ? Combine the power of R and Python 3 with Jupyter to create dynamic notebooks ? Create interactive dashboards and dynamic presentations ? Master the best coding practices and deploy your Jupyter notebooks efficiently In Detail Jupyter Notebook is a web-based environment that enables interactive computing in notebook documents. It allows you to create documents that contain live code, equations, and visualizations. This book is a comprehensive guide to getting started with data science using the popular Jupyter notebook. If you are familiar with Jupyter notebook and want to learn how to use its capabilities to perform various data science tasks, this is the book for you! From data exploration to visualization, this book will take you through every step of the way in implementing an effective data science pipeline using Jupyter. You will also see how you can utilize Jupyter's features to share your documents and codes with your colleagues. The book also explains how Python 3, R, and Julia can be integrated with Jupyter for various data science tasks. By the end of this book, you will comfortably leverage the power of Jupyter to perform various tasks in data science successfully. Style and approach This book is a perfect blend of concepts and practical examples, written in a way that is very easy to understand and implement. It follows a logical flow where you will be able to build on your understanding of the different Jupyter features with every chapter.
目录
Title Page
Copyright
Jupyter for Data Science
Credits
About the Author
About the Reviewers
www.PacktPub.com
Why subscribe?
Customer Feedback
Preface
What this book covers
What you need for this book
Who this book is for
Conventions
Reader feedback
Customer support
Downloading the example code
Errata
Piracy
Questions
Jupyter and Data Science
Jupyter concepts
A first look at the Jupyter user interface
Detailing the Jupyter tabs
What actions can I perform with Jupyter?
What objects can Jupyter manipulate?
Viewing the Jupyter project display
File menu
Edit menu
View menu
Insert menu
Cell menu
Kernel menu
Help menu
Icon toolbar
How does it look when we execute scripts?
Industry data science usage
Real life examples
Finance, Python - European call option valuation
Finance, Python - Monte Carlo pricing
Gambling, R - betting analysis
Insurance, R - non-life insurance pricing
Consumer products, R - marketing effectiveness
Using Docker with Jupyter
Using a public Docker service
Installing Docker on your machine
How to share notebooks with others
Can you email a notebook?
Sharing a notebook on Google Drive
Sharing on GitHub
Store as HTML on a web server
Install Jupyter on a web server
How can you secure a notebook?
Access control
Malicious content
Summary
Working with Analytical Data on Jupyter
Data scraping with a Python notebook
Using heavy-duty data processing functions in Jupyter
Using NumPy functions in Jupyter
Using pandas in Jupyter
Use pandas to read text files in Jupyter
Use pandas to read Excel files in Jupyter
Using pandas to work with data frames
Using the groupby function in a data frame
Manipulating columns in a data frame
Calculating outliers in a data frame
Using SciPy in Jupyter
Using SciPy integration in Jupyter
Using SciPy optimization in Jupyter
Using SciPy interpolation in Jupyter
Using SciPy Fourier Transforms in Jupyter
Using SciPy linear algebra in Jupyter
Expanding on panda data frames in Jupyter
Sorting and filtering data frames in Jupyter/IPython
Filtering a data frame
Sorting a data frame
Summary
Data Visualization and Prediction
Make a prediction using scikit-learn
Make a prediction using R
Interactive visualization
Plotting using Plotly
Creating a human density map
Draw a histogram of social data
Plotting 3D data
Summary
Data Mining and SQL Queries
Special note for Windows installation
Using Spark to analyze data
Another MapReduce example
Using SparkSession and SQL
Combining datasets
Loading JSON into Spark
Using Spark pivot
Summary
R with Jupyter
How to set up R for Jupyter
R data analysis of the 2016 US election demographics
Analyzing 2016 voter registration and voting
Analyzing changes in college admissions
Predicting airplane arrival time
Summary
Data Wrangling
Reading a CSV file
Reading another CSV file
Manipulating data with dplyr
Converting a data frame to a dplyr table
Getting a quick overview of the data value ranges
Sampling a dataset
Filtering rows in a data frame
Adding a column to a data frame
Obtaining a summary on a calculated field
Piping data between functions
Obtaining the 99% quantile
Obtaining a summary on grouped data
Tidying up data with tidyr
Summary
Jupyter Dashboards
Visualizing glyph ready data
Publishing a notebook
Font markdown
List markdown
Heading markdown
Table markdown
Code markdown
More markdown
Creating a Shiny dashboard
R application coding
Publishing your dashboard
Building standalone dashboards
Summary
Statistical Modeling
Converting JSON to CSV
Evaluating Yelp reviews
Summary data
Review spread
Finding the top rated firms
Finding the most rated firms
Finding all ratings for a top rated firm
Determining the correlation between ratings and number of reviews
Building a model of reviews
Using Python to compare ratings
Visualizing average ratings by cuisine
Arbitrary search of ratings
Determining relationships between number of ratings and ratings
Summary
Machine Learning Using Jupyter
Naive Bayes
Naive Bayes using R
Naive Bayes using Python
Nearest neighbor estimator
Nearest neighbor using R
Nearest neighbor using Python
Decision trees
Decision trees in R
Decision trees in Python
Neural networks
Neural networks in R
Random forests
Random forests in R
Summary
Optimizing Jupyter Notebooks
Deploying notebooks
Deploying to JupyterHub
Installing JupyterHub
Accessing a JupyterHub Installation
Jupyter hosting
Optimizing your script
Optimizing your Python scripts
Determining how long a script takes
Using Python regular expressions
Using Python string handling
Minimizing loop operations
Profiling your script
Optimizing your R scripts
Using microbenchmark to profile R script
Modifying provided functionality
Optimizing name lookup
Optimizing data frame value extraction
Changing R Implementation
Changing algorithms
Monitoring Jupyter
Caching your notebook
Securing a notebook
Managing notebook authorization
Securing notebook content
Scaling Jupyter Notebooks
Sharing Jupyter Notebooks
Sharing Jupyter Notebook on a notebook server
Sharing encrypted Jupyter Notebook on a notebook server
Sharing notebook on a web server
Sharing notebook on Docker
Converting a notebook
Versioning a notebook
Summary
- 神秘超市(精装)(孙诗洋)
- 离散数学简明教程(朱怀宏)
- 【13】行政纠纷裁判规则理解与适用(国家法官学院最高人民法院司法案例研究院)
- (四色)世界名人非常之路.医学世家走出的医圣:李时珍(王莲凤)
- 园林景观设计SketchUp 2014从入门到精通-(含1DVD)(麓山文化)
- 小学语文课外阅读世界文学经典名著:克雷洛夫寓言(克雷洛夫)
- 公侯世家-百变马丁-漫画《史记》故事(洋洋兔)
- 伍尔夫作品集夜与日(上下)((英)弗吉尼亚・伍尔夫)
