save
$14.01Preparing Data for Analysis: From Raw to Ready
Stop wasting 80% of your project time on data cleaning. This step-by-step guide walks you through preparing raw data for analysis with R and Stata examples.
Want to preview before you buy? Click Request Free Sample and tell us which chapter or pages you’d like to see, along with your email address. We’ll send you those pages from the eBook so you can be sure it’s the right one before buying.
- Instant download & lifetime access Available as soon as your payment is confirmed. Your download link is also sent to the email address you provide at checkout, and you can download your purchased eBook anytime from My Account → Downloads with unlimited downloads. An account is created automatically using the same email address.
- Own your eBook Receive the available PDF/EPUB file directly after purchase. No rental or temporary access — download and keep your eBook for personal use with lifetime access.
- Full refund guarantee If the eBook you receive does not match the details and description provided on the product page, contact [email protected] and we’ll issue a full refund within 24 hours.
- Secure payment through PayPal Pay securely by debit or credit card or with your PayPal account. Your card details are processed by PayPal and are never stored on our site.
$21.99$36.00
The Problem This Book Solves
Raw data rarely arrives analysis-ready. Missing values coded as 99, inconsistent string entries like “Male”/”male”/”M”, date formats that refuse to parse, and variables named in cryptic abbreviations—these are the everyday frustrations of any quantitative researcher. Cleaning this mess often consumes 80% of the project timeline, yet most statistics textbooks devote only a few pages to the topic. Bianca Manago's Preparing Data for Analysis: From Raw to Ready fills that gap with a dedicated, systematic framework for transforming messy raw data into a clean, documented dataset ready for statistical analysis.
This book is not a theory-heavy volume on advanced statistics. It is a practical, step-by-step guide to data preparation—the most time-consuming yet least-taught phase of quantitative research. Written for researchers using either R or Stata, it walks you through variable naming, data examination, and transformation with real-world examples drawn from survey data and administrative records.
About Preparing Data for Analysis: From Raw to Ready
Preparing Data for Analysis: From Raw to Ready is a concise, applied textbook published by Sage Publications in 2026 as part of the Quantitative Applications in the Social Sciences (QASS) series. The entire book focuses exclusively on the process of cleaning raw data—commonly called data cleaning or data wrangling—covering every stage from initial file compilation to final documentation.
The book follows a logical pipeline: starting with how to organize raw data files and develop a consistent variable naming scheme, then moving to thorough data examination (distributions, outliers, missing data patterns), and finally to recoding, transforming, and merging datasets. Each step is illustrated with code examples in both R and Stata, making it equally useful regardless of your software preference. The emphasis on reproducibility—including how to create a clean codebook and document every decision—sets it apart from ad hoc cleaning approaches.
Is This Book Right for You?
This book is essential for graduate students in the social sciences who are preparing their thesis data, as well as for faculty and professional researchers who regularly work with survey data, administrative records, or secondary datasets like the General Social Survey or the Panel Study of Income Dynamics. If you have ever struggled with missing data codes, inconsistent scales, or variable labels that make no sense when you return to a project months later, this guide is for you.
It is also ideal for instructors teaching courses in research methods, data science, or quantitative analysis who want to give students a dedicated resource on data preparation—a topic often glossed over in standard statistics textbooks. Program evaluators, policy analysts, and data journalists will also find the reproducible workflows directly applicable to their daily work.
Key Takeaways
- How to organize raw data files and supporting documentation for full reproducibility
- Best practices for naming variables and labeling values consistently across projects
- Techniques for examining data distributions, identifying outliers, and assessing missing data patterns
- Methods for recoding categorical variables, reversing scales, and collapsing response categories
- How to merge multiple datasets and resolve identifier mismatches, duplicate records, and ambiguous joins
- Strategies for documenting every cleaning decision in a transparent, shareable codebook
- Practical R and Stata workflows for each step of the data preparation process, with ready-to-adapt code
- How to validate cleaned data to ensure transformations were applied correctly
What Sets It Apart
Most statistics textbooks treat data cleaning as an afterthought—a few scattered pages or a single chapter at most. Manago's book dedicates an entire volume to the subject, matching the time researchers actually spend on data preparation. Unlike generic data science books that assume you already have tidy data, this one starts with raw files as they come from a survey vendor or public repository.
Competing titles like Data Wrangling with R or Data Management for Social Scientists often require prior programming experience and cover only one language. Preparing Data for Analysis begins at the very beginning—from raw files—and provides parallel code in both R and Stata, saving you from buying two resources or manually translating syntax. Its focus on social science data (surveys, administrative records) makes it more relevant than general data science texts that use marketing or tech examples.
Furthermore, the book's strong emphasis on documentation and reproducibility aligns with current best practices in open science and preregistration, giving readers a competitive edge in publishing and grant reviews.
About the Author
Bianca Manago is a social science researcher with deep expertise in data management and quantitative methods. While specific biographical details are not provided, the content reflects hands-on familiarity with the challenges of real research data—not just textbook ideals. The book's practical tone and concrete examples demonstrate that the author has worked through these problems personally.
Sage Publications is a world-leading academic publisher in the social sciences, synonymous with the QASS series (the iconic "little green books"). For decades, Sage has set the standard for concise, authoritative methods guides used by researchers at every career stage. A book carrying the Sage imprint has passed rigorous peer review and editorial vetting, providing institutional trust. The 2026 edition ensures that the content reflects current software versions and contemporary data management best practices.
Is It Worth It?
Yes—unequivocally. If you conduct quantitative research in the social sciences, data cleaning is not optional; it is a fundamental skill. This book saves you hours of guesswork and prevents costly errors that can invalidate months of analysis. It is concise enough to read in a weekend yet detailed enough to serve as a reference throughout a multi-year project.
Compared to the time and money wasted by cleaning data without a systematic method—or the reputational cost of publishing results based on flawed data—the price of this book is negligible. It is an investment in research integrity, efficiency, and the credibility of your findings.
Get Your Copy Today
Stop fighting with messy datasets. Whether you are a graduate student cleaning your dissertation survey, a faculty researcher merging multiple waves of panel data, or a policy analyst preparing administrative records for evaluation, Preparing Data for Analysis: From Raw to Ready gives you the systematic, reproducible workflow you need. Order your digital copy today from Sage, a trusted name in social science methods, and start transforming raw data into analysis-ready datasets with confidence.









