Topic outline

  • General

  • 0. Getting Started

    • 🚀 Video Introduction (9 Minutes)

      This section discusses the implications of using Python to analyze Big (qualitative) Data for social research. We also set up our working environment using a Jupyter Notebook on Google Colab.

      Keywords: .py, .ipynb, notebooks, markdown

       

      🧨 Application (10 Minutes)

      In the application, we build a simple notebook, including code and formated text (markdown)

       🥣 Ingredients:

       

      🍒 Further Food For Thought

      If you want to learn more about using Python in the digital humanities, here are some resources:

      - Dave Mattingly's excellent youTube channel Python for the Humanities features a very good introduction to the topic and includes many hands-on videos for methods and approaches in the DH.
      -Rob Mulla is a data scientist who focuses on using Python for big data. I especially recommend his introduction to using Pandas. (youtube: Rob Mulla)

    • File icon

      32907 tweets, ranging from 2010 to 2023.
      Sources:
      https://www.kaggle.com/datasets/gpreda/elon-musk-tweets
      own scraping

    • File icon

      26311 tweets, scraped from X with the following query:
      "#KAMALAHARRIS", "#DonaldTrump", "#Trump2024", "#HARRISWALZ2024"
      September 12 – November 30

       

    • File icon

      8772 tweets scraped from X on the following keywords:
      "#prochoice", "#prolife", "#mybodymychoice", "#abortion", "#familyvalues","#RightToLife","#ReproductiveRights", "#WomensHealth", "#abortiondebate", "#WomansRightToChoose"

      November 1 to November 6


  • 1. Variables and Data Types

    • 🚀 Video Introduction (6 Minutes)

      In this lection, we will take a look at the very basics of writing Python: Variables. 
      Think of a variable as a box that can store different items. What you can do with a variable largely depends on its content: Numbers strings of letter, lists.
      Keywords: variables, len(), split()

       

      🧨 Application (10 Minutes)

      In the application, we build a simple pipeline to clean tweets by using the lower() function.

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to learn more about variables and how to apply them.
      You find the notebook here:

      The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.

    • 🚀 Video Introduction (9 Minutes)

      This section discusses the implications of using Python to analyze Big (qualitative) Data for social research. We also set up our working environment using a Jupyter Notebook on Google Colab.

      Keywords: .py, .ipynb, notebooks, markdown

       

      🧨 Application (10 Minutes)

      In the application, we build a simple notebook, including code and formated text (markdown)

       🥣 Ingredients:

       

      🍒 Further Food For Thought

      If you want to learn more about using Python in the digital humanities, here are some resources:

      - Dave Mattingly's excellent youTube channel Python for the Humanities features a very good introduction to the topic and includes many hands-on videos for methods and approaches in the DH.
      -Rob Mulla is a data scientist who focuses on using Python for big data. I especially recommend his introduction to using Pandas. (youtube: Rob Mulla)

  • 2 – Fun with Lists

    • 🚀 Video Introduction (9 Minutes)

      If a variable is a box, the list can be thought of as a shelf with room for storage. Lists are fun to use, as everything that works in a small format can be scaled up to massive datasets. I'll show you how sentences can be split into lists for further options.

      keywords: list, split(), len()

       

      🧨 Application (9 Minutes)

      Further things with lists. Also: Have you ever wondered if the first chapter of Moby Dick contains the word 'watery'? Let's find out …

       
       🥣 Ingredients:
      - Moby Dick from Gutenberg

       

       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to learn more about variables and how to apply them.  Use this interactive notebook to deepen your knowledge:
      Unit 2: Lists (The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

    • 🚀 Video Introduction (6 Minutes)

      In this lection, we will take a look at the very basics of writing Python: Variables. 
      Think of a variable as a box that can store different items. What you can do with a variable largely depends on its content: Numbers strings of letter, lists.
      Keywords: variables, len(), split()

       

      🧨 Application (10 Minutes)

      In the application, we build a simple pipeline to clean tweets by using the lower() function.

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to learn more about variables and how to apply them.
      You find the notebook here:

      The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.

  • 3 – Stuck in a Loop

    • 🚀 Video Introduction (10 Minutes)

      Everything you do more than twice – automatize! With the for loop you can repeat an operation on every item in a list. Want to to turn 10.000 tweets into lowercase? No problem …

      keywords: for item in, lower()

       

      🧨 Application (11 Minutes)

      We build a simple pipeline to make tweets better searchable and remove special characters for further analysis. And: There are squirrels …
       
    • 🚀 Video Introduction (9 Minutes)

      If a variable is a box, the list can be thought of as a shelf with room for storage. Lists are fun to use, as everything that works in a small format can be scaled up to massive datasets. I'll show you how sentences can be split into lists for further options.

      keywords: list, split(), len()

       

      🧨 Application (9 Minutes)

      Further things with lists. Also: Have you ever wondered if the first chapter of Moby Dick contains the word 'watery'? Let's find out …

       
       🥣 Ingredients:
      - Moby Dick from Gutenberg

       

       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to learn more about variables and how to apply them.  Use this interactive notebook to deepen your knowledge:
      Unit 2: Lists (The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

  • 4 – Scaling Things up with Modules

    • 🚀 Video Introduction (7 Minutes)

      Python draws on a plethora of open-source programs – called modules. From giant sets of trained data to the transformer model that powers Google search, all of these are available with some lines of code. In this section, we import modules into our code and scale up what's possible.

      keywords: import, pip install, dir()

       

      🧨 Application (10 Minutes)

      We import KeyBert, a powerful library that recognizes important words in a text. KeyBert draws on Google's transformer BERT and is your first step toward topic modeling.
       
       🥣 Ingredients:
      - The KeyBert Module by Maarten Grootendorst
      - A random paper, although you can take this one
       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to learn more of KeyBert's parameters. Use this interactive notebook to deepen your knowledge:
      Unit 4: Modules (The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

    • 🚀 Video Introduction (10 Minutes)

      Everything you do more than twice – automatize! With the for loop you can repeat an operation on every item in a list. Want to to turn 10.000 tweets into lowercase? No problem …

      keywords: for item in, lower()

       

      🧨 Application (11 Minutes)

      We build a simple pipeline to make tweets better searchable and remove special characters for further analysis. And: There are squirrels …
       

       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to learn more about variables and how to apply them.  Use this interactive notebook to deepen your knowledge:
      Unit 3: For-Loops (The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

       
  • 5 – Read and Write Text

    • 🚀 Video Introduction (11 Minutes)

      Textfiles are a simple and versatile way to access your data and export them to other applications. With the open()-function you can both read and write texts in Python.
       

      keywords: with open(), write(), read()

       

      🧨 Application (11 Minutes)

      We import KeyBert, a powerful library that recognizes important words in a text. KeyBert draws on Google's transformer BERT and is your first step toward topic modeling.
       
       
       🥣 Ingredients:
      - Comments from Reddit
       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to explore a little project with comments from Reddit that we import and analyze:
      Unit 5: Working with Text. The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

    • 🚀 Video Introduction (7 Minutes)

      Python draws on a plethora of open-source programs – called modules. From giant sets of trained data to the transformer model that powers Google search, all of these are available with some lines of code. In this section, we import modules into our code and scale up what's possible.

      keywords: import, pip install, dir()

       

      🧨 Application (10 Minutes)

      We import KeyBert, a powerful library that recognizes important words in a text. KeyBert draws on Google's transformer BERT and is your first step toward topic modeling.
       
       🥣 Ingredients:
      - The KeyBert Module by Maarten Grootendorst
      - A random paper, although you can take this one
       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to learn more of KeyBert's parameters. Use this interactive notebook to deepen your knowledge:
      Unit 4: Modules (The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

  • Unit 6 IF THIS THEN WTF?!

    • 🚀 Video Introduction (11 Minutes)

      Textfiles are a simple and versatile way to access your data and export them to other applications. With the open()-function you can both read and write texts in Python.
       

      keywords: with open(), write(), read()

       

      🧨 Application (11 Minutes)

      We import KeyBert, a powerful library that recognizes important words in a text. KeyBert draws on Google's transformer BERT and is your first step toward topic modeling.
       
       
       🥣 Ingredients:
      - Comments from Reddit
       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to explore a little project with comments from Reddit that we import and analyze:
      Unit 5: Working with Text. The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

    • 🚀 Video Introduction (11 Minutes)

      Textfiles are a simple and versatile way to access your data and export them to other applications. With the open()-function you can both read and write texts in Python.
       

      keywords: with open(), write(), read()

       

      🧨 Application (11 Minutes)

      We import KeyBert, a powerful library that recognizes important words in a text. KeyBert draws on Google's transformer BERT and is your first step toward topic modeling.
       
       
       🥣 Ingredients:
      - Comments from Reddit
       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to explore a little project with comments from Reddit that we import and analyze:
      Unit 5: Working with Text. The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

  • Unit 7 – Dictionaries & Sets

    • 🚀 Video Introduction (11 Minutes)

      In some cases, you need to store data that is more structured. While lists are 'flat', which means that all items appear on the same level, dictionaries offer you the possibility to create dimensions (keys) where data (values) can be stored. Think of it like the head and body of a table. In Python, dictionaries, sets and tuples bring new options to the table of data cleaning.

      keywords: dict{}, set{}, 

       

      🧨 Application (11 Minutes)

      We read a list of hashtags into Python and try to make sense of it by filtering out the unique values.
       🥣 Ingredients:
      - A random list of hashtags: hashtags.txt
       

       

    • 🚀 Video Introduction (11 Minutes)

      Conditionals help you to create more complex programs, taking a different direction depending on the outcome. You can use IF / ELSE to filter data and create sophisticated search queries.
       

      keywords: if, elif, else

        

       

      🧨 Application (11 Minutes)

      We use conditionals to filter tweets.
       
       
       

      🍒 Further Food For Thought

      You want to know more? Deep-dive into this notebook to explore more about lists and boolean statements. In section 3.2 we build a small chatbot that takes user input to assess medical symptoms. The notebook is protected, but you can create a copy of it in your own Colab repository by clicking File – Save Copy in Drive.)

  • Unit 8 – Make Things Do Things with Functions

    • 🚀 Video Introduction (9 Minutes)

      there might probably be a life in Python without creating your own functions – but it is stale and unfortunate. Think of functions as little (or massive) tools you build for every task appearing repeatedly in your code. In this video I show you how to create functions and use them in your code.
       
      keywords: def my_function()

       

      🧨 Application (11 Minutes)

      We use the library flair to read sentiments from our data and build that into a function.
       🥣 Ingredients:
      - some random tweets
      - the powerful library flair (use pip install flair)
       
    • 🚀 Video Introduction (11 Minutes)

      In some cases, you need to store data that is more structured. While lists are 'flat', which means that all items appear on the same level, dictionaries offer you the possibility to create dimensions (keys) where data (values) can be stored. Think of it like the head and body of a table. In Python, dictionaries, sets and tuples bring new options to the table of data cleaning.

      keywords: dict{}, set{}, 

       

      🧨 Application (11 Minutes)

      We read a list of hashtags into Python and try to make sense of it by filtering out the unique values.
       🥣 Ingredients:
      - A random list of hashtags: hashtags.txt
       

       

  • Unit 9 – Wrangling Data With Pandas

    • 🚀 Video Introduction (5 Minutes)

      Today, we meet another animal from the Python zoo: Pandas (which stands for panel data) is the most essential library to tackle data, and definitely worth checking out. Pandas is a sophisticated toolset allowing you to read data into a databank structure (called data frame), search through it, and perform (simple) analyses and visualizations. As most of the work with Python includes data (text/numbers/metadata), Pandas is the backbone of many projects. 
       
      keywords: import pandas as pd

       

      🧨 Application (19 Minutes)

      We use pandas to read a dataset of Elon Musk's tweets into a data frame.
       🥣 Ingredients:
      - some random tweets by Elon Musk, if you wat to get the full dataset, see here
      - the powerful library pandas (use pip install pandas)
       

      🍒 Further Food For Thought

      If you want to learn more about pandas, I recommend the excellent videos by data scientist Rob Mulla. You can start with his excellent introduction to Pandas.

    • 🚀 Video Introduction (9 Minutes)

      there might probably be a life in Python without creating your own functions – but it is stale and unfortunate. Think of functions as little (or massive) tools you build for every task appearing repeatedly in your code. In this video I show you how to create functions and use them in your code.
       
      keywords: def my_function()

       

      🧨 Application (11 Minutes)

      We use the library flair to read sentiments from our data and build that into a function.
       🥣 Ingredients:
      - some random tweets
      - the powerful library flair (use pip install flair)
       
  • Unit 10 – Topic Modeling

    • 🚀 Video Introduction (5 Minutes)

      Topic Modeling has become one of the most essential methods of analyzing big qualitative data in social and political sciences (Isoaho et al., 2021, p. 2). In this course, I introduce it as an exploratory first step to distant-read texts. Although it draws on complex statistical operations under the hood, performing a TM in Python is straight forward and only needs a few lines of code.
       
      keywords: topic, k, bert
       

      🧨 Project: Topic Modeling (25 Minutes)

      We use TM to distant-read through 8.000 tweets on the US elections in 2024.
       🥣 Ingredients:
      - as data set scraped from X in October 2024.
      - the powerful library BERTopic (use pip install bertopic)
       

      🍒 Further Food For Thought

      If you are interested in introducing TM in a research paper, here're some sources to start from:

    • 🚀 Video Introduction (5 Minutes)

      Today, we meet another animal from the Python zoo: Pandas (which stands for panel data) is the most essential library to tackle data, and definitely worth checking out. Pandas is a sophisticated toolset allowing you to read data into a databank structure (called data frame), search through it, and perform (simple) analyses and visualizations. As most of the work with Python includes data (text/numbers/metadata), Pandas is the backbone of many projects. 
       
      keywords: import pandas as pd

       

      🧨 Application (19 Minutes)

      We use pandas to read a dataset of Elon Musk's tweets into a data frame.
       🥣 Ingredients:
      - some random tweets by Elon Musk, if you wat to get the full dataset, see here
      - the powerful library pandas (use pip install pandas)
       

      🍒 Further Food For Thought

      Wonder how a hashtag appears over time? Here's a working Notebook that creates a timeline and performs other data queries.

      If you want to learn more about pandas, I recommend the excellent videos by data scientist Rob Mulla. You can start with his excellent introduction to Pandas.

  • Unit 11 – Project: Sentiment Analysis

  • Unit 12 – Project: Scraping Reddit

    • 🧨 Project: Scraping Reddit 

      Reddit is a complex microcosm hosting a trillion of public subspheres, which makes scraping Reddit interesting for the social researcher. In this project, we will use the powerful library PRAW to get data from Reddit APIs; before you start, please make sure to follow this short tutorial below to create your own developer account. Also, please make sure to check  PRAWs documentation for further options. 
       
      You can either use this preconfigured notebook, which can be modified according to your needs, or follow my step-by-step instructions.
       
      Part 1: Introduction, Scraping Single Comments (25 minutes)
      Part 2: Building a Scraper, Scraping Entire Queries (22 minutes)

       

    •  🥣 Ingredients: Setting up a Reddit Account

      PRAW uses Reddit’s API – which you can think of as a backdoor to Reddit’s database. To use the library you first have to get a Reddit account and then create an ‚app‘ for scraping. Please follow these instructions:

      1. Go to reddit.com and register as a user
        2. Once you’re logged in, please see this site, which is Reddits developers section.
      2. Click on ‚create new app’ 

      1.  make sure to choose ’script and provide the ‚name‘ and ‚developer. Redirect URL can be any url or left bank.
        3. After finishing you can copy the following information to your Python script




      


    • Difficulty: ️🌶️
      Based on: Lists, For-Loops, Dictionaries
      Working Notebook

       

    • 🧨 Project: Renaming Papers from Sage (28 +25 Minutes)

      This project is based on a program I use on a daily basis when researching literature. It take a messy filename like mcdonnell-et-al-2023-this-is-your-brain-on-autopilot-2-0-the-influence-of-practice-on-driver-workload-and-engagement and turns into something readable like McDonnell (2023). This Is Your Brain on Autopilot 2.0 - The Influence of Practice on Driver Workload and Engagement During On-Road.pdf. In this project you get to know three powerful  libraries: pathlib to work with files on your computer, pyPDF2 to read pdfs and habanero to look up scholarly metadata on crossref.
       
      keywords: path, dicitionaries, lists, pdf, regex

      Video 1 (Read Files, read PDFs, get DOIs)

      Video 2 (Get Metadata, write files)

       🥣 Ingredients:
      - the libraries habanero, pypdf2 and pathlib (use pip install)
       
  • Annotated Projects

    • Scraping Data on German Party Donations with Beautiful Soup

       

      Here's a fully working project from the course ’Data Driven Journalism’. The German Parliament (Deutscher Bundestag) publishes party donations higher than 50.000€ – the donations from the last 30 years can be found online. However, the lists are very messy so scraping and cleaning them is crucial. This annotated project goes through every step with clear comments. If you execute each cell, you will get a csv-dataset containing all donations from 2010 — 2023.

    • Besides analyzing data and handling databases, Python has powerful modules to create data visualizations. Today we will take a look at seaborn and bokeh, the latter providing interactive DataViz that could be embedded in Websites. 

      https://colab.research.google.com/drive/1Aowbvaa00ArhJPdFP7pqFzz1NafDj3U1?usp=sharing

       

  • Useful Resources

    • Videos and Online-Courses

      A great resource for valuable video courses is learning.oreilly.com – which can be accessed for free with the h_da single-sign-on. I especially recommend the videos by Arianne Dee, who is a very thoughtful instructor. I did her course ‘Next Level Python’ but I guess the “Introduction to Python: Learn How to Program Today with Python” is also worth taking a look. An advanced course on using Python for data visualization is this one by AI Sciences.

      Udemy is a commercial platform offering a lot of courses on Python. I definitely recommend the introduction by my former colleague René Brunner (in German Language), which was actually the first course I enrolled for. It has a slow pace and is very comprehensive. If you like I can ask him for a discount.


      A great course for understanding the advanced concept of Sentiment Analysis is
      Applied Text Mining and Sentiment Analysis with Python by Data Analyst Benjamin Termonia. He not only offers a great introduction to the concept but also a step-by-step tutorial on building/training your own Sentiment Analysis environment.

       

      Code Academy is a very innovative commercial platform that is totally worth its high price of $149.99/year for students as it provides clearly structured lessons and learning paths for different fields, an interactive online learning editor, videos, cheat sheets etc. It’s definitely worth applying for the free trial period. 


      Finally, the resource python for humanities offers great tutorials on advanced methods like sentiment analysis and network graphs. This is also your first address when it comes to using Python in a scientific context. The author is a fellow at the Smithsonian and offers great insights into methods like ML, Topic Modeling, Sentiment Analysis etc.

      Other Ressources

      There are plenty of sites to look up basic functions and test-run Python's core concepts. For starters, I recommend w3-schools.

      If you have concrete questions about your code, StackOverflow is the place to go. If you are searching for discussion around libraries, check GitHub.

       

       

  • Data Repository

    • File icon
    • Scraping Data on German Party Donations with Beautiful Soup

       

      Here's a fully working project from the course ’Data Driven Journalism’. The German Parliament (Deutscher Bundestag) publishes party donations higher than 50.000€ – the donations from the last 30 years can be found online. However, the lists are very messy so scraping and cleaning them is crucial. This annotated project goes through every step with clear comments. If you execute each cell, you will get a csv-dataset containing all donations from 2010 — 2023.