diff --git a/episodes/conditionals.md b/episodes/conditionals.md index 1616e805..ee656f23 100644 --- a/episodes/conditionals.md +++ b/episodes/conditionals.md @@ -76,6 +76,52 @@ for checkout in checkouts: Notice that our `else` statement led to a false output that says 10 is under the limit. We can address this by adding a different kind of `else` statement. +::::::::::::::::::::::::::::::::::::::: challenge + +## Age conditionals + +Write a Python program that checks the age of a user to determine if they will receive a youth or adult library card. The program should: + +1. Store `age` in a variable. +2. Use an `if` statement to check if the age is 16 or older. If true, print "You are eligible for an adult library card." +3. Use an `else` statement to print "You are eligible for a youth library card" if the age is less than 16. + +If you finish early, try this challenge: + +- In a new cell, adapt your program to loop through a list of age values, testing each age with the same output as above. + +::::::::::::::: solution + +## Solution + +For parts 1 to 3: + +```python +age = 25 + +if age >= 16: + print('You are eligible for an adult library card.') +else: + print('You are eligible for a youth library card.') +``` + +For the challenge: +```python +ages = [10, 16, 30, 65] + +for age in ages: + if age >= 16: + print('You are eligible for an adult library card.') + else: + print('You are eligible for a youth library card.') +``` + + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + + ## Use `elif` to specify additional tests. You can use `elif` (short for "else if") to provide several alternative choices, each with its own test. An `elif` statement should always be associated with an `if` statement, and must come before the `else` statement (which is the catch all). @@ -153,52 +199,6 @@ for user in users: *Warning*: 120 is over the grad limit. ``` -::::::::::::::::::::::::::::::::::::::: challenge - -## Age conditionals - -Write a Python program that checks the age of a user to determine if they will receive a youth or adult library card. The program should: - -1. Store `age` in a variable. -2. Use an `if` statement to check if the age is 16 or older. If true, print "You are eligible for an adult library card." -3. Use an `else` statement to print "You are eligible for a youth library card" if the age is less than 16. - -If you finish early, try this challenge: - -- In a new cell, adapt your program to loop through a list of age values, testing each age with the same output as above. - -::::::::::::::: solution - -## Solution - -For parts 1 to 3: - -```python -age = 25 - -if age >= 16: - print('You are eligible for an adult library card.') -else: - print('You are eligible for a youth library card.') -``` - -For the challenge: -```python -ages = [10, 16, 30, 65] - -for age in ages: - if age >= 16: - print('You are eligible for an adult library card.') - else: - print('You are eligible for a youth library card.') -``` - - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: - - ::::::::::::::::::::::::::::::::::::::: challenge ## Conditional logic: Fill in the blanks diff --git a/episodes/data-visualisation.md b/episodes/data-visualisation.md index 0b834a02..f2a9accf 100644 --- a/episodes/data-visualisation.md +++ b/episodes/data-visualisation.md @@ -138,55 +138,6 @@ albany['circulation'].plot(kind='hist', bins=20, {alt="histogram of the Albany branch circulation."} -## Use Plotly for interactive plots - -Let’s switch back to the full DataFrame in `df_long` and use another -plotting package in Python called Plotly. - -```python -import plotly.express as px -``` - -Now we can visualize how circulation counts have changed over time for selected branches. This can be especially useful for identifying trends, seasonality, or data anomalies. We willfirst create a subset of our data to look at branches starting with the letter 'A'. Feel free to select different branches. After subsetting, we will sort our new DataFrame by date and then plot our data by date and circulation count. - -``` python -# Creating a line plot for a few selected branches to avoid clutter -selected_branches = df_long[df_long['branch'].isin(['Altgeld', - 'Archer Heights', - 'Austin', - 'Austin-Irving', - 'Avalon'])] -selected_branches = selected_branches.sort_values(by='date') -``` - -``` python -fig = px.line(selected_branches, x=selected_branches.index, y='circulation', color='branch', title='Circulation Over Time for Selected Branches') -fig.show() -``` - -Here is a view of the [interactive output of the Plotly line chart](learners/line_plot_int.html). - - -One advantage that Plotly provides over Matplotlib is that it has some interactive features out of the box. Hover your cursor over the lines in the output to find out more granular data about specific branches over time. - - -### Bar plots with Plotly - -Let’s use a barplot to compare the distribution of circulation counts -among branches. We first need to group our data by branch and sum up the circulation counts. Then we can use the bar plot to show the -distribution of total circulation over branches. - -``` python -# Aggregate circulation by branch -total_circulation_by_branch = df_long.groupby('branch')['circulation'].sum().reset_index() - -# Create a bar plot -fig = px.bar(total_circulation_by_branch, x='branch', y='circulation', title='Total Circulation by Branch') -fig.show() -``` - -Here is a view of the [interactive output of the Plotly bar chart](learners/bar_plot_int.html). - ::::::::::::::::::::::::::::::::::::::: challenge ## Plotting with Pandas @@ -248,6 +199,56 @@ uptown['circulation'].plot(title='Uptown Circulation', :::::::::::::::::::::::::::::::::::::::::::::::::: + +## Use Plotly for interactive plots + +Let’s switch back to the full DataFrame in `df_long` and use another +plotting package in Python called Plotly. + +```python +import plotly.express as px +``` + +Now we can visualize how circulation counts have changed over time for selected branches. This can be especially useful for identifying trends, seasonality, or data anomalies. We willfirst create a subset of our data to look at branches starting with the letter 'A'. Feel free to select different branches. After subsetting, we will sort our new DataFrame by date and then plot our data by date and circulation count. + +``` python +# Creating a line plot for a few selected branches to avoid clutter +selected_branches = df_long[df_long['branch'].isin(['Altgeld', + 'Archer Heights', + 'Austin', + 'Austin-Irving', + 'Avalon'])] +selected_branches = selected_branches.sort_values(by='date') +``` + +``` python +fig = px.line(selected_branches, x=selected_branches.index, y='circulation', color='branch', title='Circulation Over Time for Selected Branches') +fig.show() +``` + +Here is a view of the [interactive output of the Plotly line chart](learners/line_plot_int.html). + + +One advantage that Plotly provides over Matplotlib is that it has some interactive features out of the box. Hover your cursor over the lines in the output to find out more granular data about specific branches over time. + + +### Bar plots with Plotly + +Let’s use a barplot to compare the distribution of circulation counts +among branches. We first need to group our data by branch and sum up the circulation counts. Then we can use the bar plot to show the +distribution of total circulation over branches. + +``` python +# Aggregate circulation by branch +total_circulation_by_branch = df_long.groupby('branch')['circulation'].sum().reset_index() + +# Create a bar plot +fig = px.bar(total_circulation_by_branch, x='branch', y='circulation', title='Total Circulation by Branch') +fig.show() +``` + +Here is a view of the [interactive output of the Plotly bar chart](learners/bar_plot_int.html). + ::::::::::::::::::::::::::::::::::::::: challenge ## Plot the top five branches diff --git a/episodes/for-loops.md b/episodes/for-loops.md index 5313fbb8..c9fd5279 100644 --- a/episodes/for-loops.md +++ b/episodes/for-loops.md @@ -163,6 +163,38 @@ for number in range(0,3): 2 ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## Use range() in a loop + +Print out the numbers 10, 11, 12, 13, 14, 15, using range() in a `for` loop. + +::::::::::::::: solution + +## Solution + +```python + +for num in range(10, 16): + print(num) + +``` + +```output +10 +11 +12 +13 +14 +15 +``` + + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + + ## Accumulators A common loop pattern is to initialize an *accumulator* variable to zero, an empty string, or an empty list before the loop begins. Then the loop updates the accumulator variable with values from a collection. @@ -249,37 +281,6 @@ for veg in vegetables: :::::::::::::::::::::::::::::::::::::::::::::::::: ::::::::::::::::::::::::::::::::::::::: challenge -## Use range() in a loop - -Print out the numbers 10, 11, 12, 13, 14, 15, using range() in a `for` loop. - -::::::::::::::: solution - -## Solution - -```python - -for num in range(10, 16): - print(num) - -``` - -```output -10 -11 -12 -13 -14 -15 -``` - - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: - -::::::::::::::::::::::::::::::::::::::: challenge - ## Use a string index in a loop How would you loop through a list with the values 'red', 'green', and 'blue' to create the acronym `rgb`, pulling from the first letters in each string? Print the acronym when the loop is finished. diff --git a/episodes/getting-started.md b/episodes/getting-started.md index bfff0174..2aa17a0b 100644 --- a/episodes/getting-started.md +++ b/episodes/getting-started.md @@ -18,7 +18,7 @@ exercises: 5 - How can I identify and use key features of JupyterLab to create and manage a Python notebook? - How do I run Python code in JupyterLab, and how can I see and interpret the results? -- + :::::::::::::::::::::::::::::::::::::::::::::::::: ## Why Python? @@ -151,7 +151,7 @@ If you move your cursor back to the first cell, just after the `7 * 3` code, and ```python 7 * 3 -2 +1 +2 + 1 ``` While Python runs both calculations Juypter will only display the output from the last line of code in a specific cell, unless you tell it to do otherwise. diff --git a/episodes/libraries.md b/episodes/libraries.md index da7bceab..05da8950 100644 --- a/episodes/libraries.md +++ b/episodes/libraries.md @@ -155,6 +155,51 @@ Many popular libraries have common aliases. For example: Using these common aliases can make it easier to work with existing documentation and tutorials. +::::::::::::::::::::::::::::::::::::::: challenge + +## Importing With Aliases + +1. Fill in the blanks so that the program below prints `0123456789`. +2. Rewrite the program so that it uses `import` *without* `as`. +3. Which form do you find easier to read? + +```python +import string as s +numbers = ____.digits +print(____) +``` + +::::::::::::::: solution + +## Solution + +```python +import string as s +numbers = s.digits +print(numbers) +``` + +can be written as + +```python +import string +numbers = string.digits +print(numbers) +``` + +Since you just wrote the code and are familiar with it, you might actually +find the first version easier to read. But when trying to read a huge piece +of code written by someone else, or when getting back to your own huge piece +of code after several months, non-abbreviated names are often easier, expect +where there are clear abbreviation conventions. + + + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + + ## Pandas `Pandas` is a widely-used Python library for statistics using tabular data. Essentially, it gives you access to 2-dimensional tables whose columns have names and can have different data types. We can start using pandas by reading a `Comma Separated Values` (CSV) data file with the function `pd.read_csv()`. The function `.read_csv()` expects as an argument the path to and name of the file to be read. This returns a dataframe that you can assign to a variable. @@ -321,50 +366,6 @@ This gives us, for example, the count, minimum, maximum, and mean values from ea ::::::::::::::::::::::::::::::::::::::: challenge -## Importing With Aliases - -1. Fill in the blanks so that the program below prints `0123456789`. -2. Rewrite the program so that it uses `import` *without* `as`. -3. Which form do you find easier to read? - -```python -import string as s -numbers = ____.digits -print(____) -``` - -::::::::::::::: solution - -## Solution - -```python -import string as s -numbers = s.digits -print(numbers) -``` - -can be written as - -```python -import string -numbers = string.digits -print(numbers) -``` - -Since you just wrote the code and are familiar with it, you might actually -find the first version easier to read. But when trying to read a huge piece -of code written by someone else, or when getting back to your own huge piece -of code after several months, non-abbreviated names are often easier, expect -where there are clear abbreviation conventions. - - - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: - -::::::::::::::::::::::::::::::::::::::: challenge - ## Locating the Right Module Given the variables `year`, `month` and `day`, how would you generate a date in the standard iso format: diff --git a/episodes/lists.md b/episodes/lists.md index 1c3a5faf..425fed38 100644 --- a/episodes/lists.md +++ b/episodes/lists.md @@ -58,6 +58,61 @@ First item: marc The first three items: ['marc', 'frbr', 'mets'] ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## Lists: Length and Indexing +1. Create a list named `colors` containing the strings 'red', 'blue', and 'green'. +2. Print the length of the list. +3. Print the first color using indexing. + +::::::::::::::: solution + +## Solution +```python +colors = ['red', 'blue', 'green'] +print(len(colors)) +print(colors[0]) +``` + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::: challenge + +## List slicing +1. Create a list of numbers defined as [1, 2, 3, 4, 5, 6]. +2. Print the first three items in the list using slicing. +3. Print the last three items using slicing. + +::::::::::::::: solution + +## Solution +```python +numbers = [1, 2, 3, 4, 5, 6] +print(numbers[0:3]) +print(numbers[3:6]) +``` +```output +[1, 2, 3] +[4, 5, 6] +``` + +You can also leave the first and last elements in a slice blank to refer to the first and last elements in a list: + +```python +print(numbers[:3]) +print(numbers[3:]) +``` +```output +[1, 2, 3] +[4, 5, 6] +``` + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + ## Reassign list values with their index. Use an index value along with your list variable to replace a value from the list. @@ -172,60 +227,6 @@ animals before: ['dog', 'bird', 'shark', 'dog'] animals after: ['bird', 'shark', 'dog'] ``` -::::::::::::::::::::::::::::::::::::::: challenge - -## Lists: Length and Indexing -1. Create a list named `colors` containing the strings 'red', 'blue', and 'green'. -2. Print the length of the list. -3. Print the first color using indexing. - -::::::::::::::: solution - -## Solution -```python -colors = ['red', 'blue', 'green'] -print(len(colors)) -print(colors[0]) -``` - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::: challenge - -## List slicing -1. Create a list of numbers defined as [1, 2, 3, 4, 5, 6]. -2. Print the first three items in the list using slicing. -3. Print the last three items using slicing. - -::::::::::::::: solution - -## Solution -```python -numbers = [1, 2, 3, 4, 5, 6] -print(numbers[0:3]) -print(numbers[3:6]) -``` -```output -[1, 2, 3] -[4, 5, 6] -``` - -You can also leave the first and last elements in a slice blank to refer to the first and last elements in a list: - -```python -print(numbers[:3]) -print(numbers[3:]) -``` -```output -[1, 2, 3] -[4, 5, 6] -``` - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: ::::::::::::::::::::::::::::::::::::::: challenge diff --git a/episodes/looping-data-sets.md b/episodes/looping-data-sets.md index 8b9665fc..5e9170b7 100644 --- a/episodes/looping-data-sets.md +++ b/episodes/looping-data-sets.md @@ -54,6 +54,27 @@ print(f"all csv files in data directory: {glob.glob('data/*.csv')}") all csv files in data directory: ['data/2011_circ.csv', 'data/2016_circ.csv', 'data/2017_circ.csv', 'data/2022_circ.csv', 'data/2018_circ.csv', 'data/2019_circ.csv', 'data/2012_circ.csv', 'data/2013_circ.csv', 'data/2021_circ.csv', 'data/2020_circ.csv', 'data/2015_circ.csv', 'data/2014_circ.csv'] ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## Determining Matches + +Which of these files would be matched by the expression `glob.glob('data/*circ.csv')`? + +1. `data/2011_circ.csv` +2. `data/2012_circ_stats.csv` +3. `circ/2013_circ.csv` +4. Both 1 and 3 + +::::::::::::::: solution + +## Solution + +Only item 1 is matched by the wildcard expression `data/*circ.csv`. + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + ## Use `glob` and `for` to process batches of files. Now we can use glob in a `for` loop to create DataFrames from all of the CSV files in the `data` directory. To use tools like `glob` it helps if files are named and stored consistently so that simple patterns will find the right data. You can learn more about how to name files to improve machine-readability from the [Open Science Foundation article on file naming](https://help.osf.io/article/146-file-naming). @@ -102,6 +123,37 @@ data/2021_circ.csv 271811 data/2022_circ.csv 301340 ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## Minimum circulation per year + +Modify the following code to print out the lowest value in the `ytd` column from each year/file. + +```python +import pandas as pd +for csv in sorted(glob.glob('data/*.csv')): + data = pd.read_csv(____) + print(csv, data['____'].____()) + +``` + +::::::::::::::: solution + +## Solution + +```python +import pandas as pd +for csv in sorted(glob.glob('data/*.csv')): + data = pd.read_csv(csv) + print(csv, data['ytd'].min()) + +``` + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + + ## Appending DataFrames to a list @@ -192,56 +244,6 @@ f'Number of rows in df: {len(df)}' 'Number of rows in df: 963' ``` -::::::::::::::::::::::::::::::::::::::: challenge - -## Determining Matches - -Which of these files would be matched by the expression `glob.glob('data/*circ.csv')`? - -1. `data/2011_circ.csv` -2. `data/2012_circ_stats.csv` -3. `circ/2013_circ.csv` -4. Both 1 and 3 - -::::::::::::::: solution - -## Solution - -Only item 1 is matched by the wildcard expression `data/*circ.csv`. - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: - -::::::::::::::::::::::::::::::::::::::: challenge - -## Minimum circulation per year - -Modify the following code to print out the lowest value in the `ytd` column from each year/file. - -```python -import pandas as pd -for csv in sorted(glob.glob('data/*.csv')): - data = pd.read_csv(____) - print(csv, data['____'].____()) - -``` - -::::::::::::::: solution - -## Solution - -```python -import pandas as pd -for csv in sorted(glob.glob('data/*.csv')): - data = pd.read_csv(csv) - print(csv, data['ytd'].min()) - -``` - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: ::::::::::::::::::::::::::::::::::::::: challenge diff --git a/episodes/pandas.md b/episodes/pandas.md index e73f7829..2035aa3b 100644 --- a/episodes/pandas.md +++ b/episodes/pandas.md @@ -129,6 +129,38 @@ type(df['year']) pandas.core.series.Series ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## Displaying rows and columns + +How would you use slicing and column names to select the following subsets of rows and columns from the circulation DataFrame? + +1. The city column. +2. Rows 10 to 20. +3. Rows 20 to 30 from the zip code column. + +::::::::::::::: solution + +## Solution + +```python +#1 +df['city'] + +#2 +df[10:21] + +#3 +df['zip code'][20:31] + +``` + + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + + ## Summary statistics on columns A pandas Series is a one-dimensional array, like a column in a spreadsheet, while a pandas DataFrame is a two-dimensional tabular data structure with labeled axes, similar to a spreadsheet. One of the advantages of pandas is that we can use built-in functions like `max()`, `min()`, `mean()`, and `sum()` to provide summary statistics across Series such as columns. Since it can be difficult to get a sense of the range of data in a large DataFrame by looking over the whole thing manually, these functions can help us understand our dataset quickly and ask specific questions. @@ -256,6 +288,59 @@ year branch Name: ytd, dtype: int64 ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## Unique items + +How would you display: + +1. all of the unique zip codes in the dataset? +2. the number of unique zip codes in the dataset? + + +::::::::::::::: solution + +## Solution + +```python + +#1 +df['zip code'].unique() + +#2 +df['zip code'].nunique() + +``` +::::::::::::::::::::::::: +:::::::::::::::::::::::::::::::::::::::::::::::::: + +::::::::::::::::::::::::::::::::::::::: challenge + +## Summary statistics and groupby() + +We can apply `mean()` to pandas series' in the same way we used `sum()`, `min()`, and `max()` above. How would you display the following? + +1. the mean number of ytd checkouts grouped by zip code? +2. the mean number of ytd checkouts grouped by zip code, and sorted from smallest to largest? + + +::::::::::::::: solution + +## Solution + +```python +#1 +df.groupby('zip code')['ytd'].mean() + +#2 +df.groupby('zip code')['ytd'].mean().sort_values() + +``` +::::::::::::::::::::::::: +:::::::::::::::::::::::::::::::::::::::::::::::::: + + + ## Use .iloc[] and .loc[] to select DataFrame locations. You can point to specific locations in a DataFrame using two-dimensional numerical indexes with `.iloc[]`. @@ -284,6 +369,25 @@ Branch: Albany Park YTD circ: 120059 ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## Using loc() + +How would you use `loc()` to select rows 20 to 30 from the zip code column (the same rows as the last example in the challenge above)? + +Tip: slices use "non-inclusive" indexing -- so require you to ask for `df[10:21]` to see row 20, but `loc()` uses inclusive indexing. + +::::::::::::::: solution + +## Solution + +```python +df.loc[20:30, 'zip code'] + +``` +::::::::::::::::::::::::: +:::::::::::::::::::::::::::::::::::::::::::::::::: + ## Save DataFrames You might want to export the series of usage by year and branch that we just created so that you can share it with colleagues. Pandas includes a variety of methods that begin with `.to_...` that allow us to convert and export data in different ways. First, let's save our series as a DataFrame so we can view the output in a better format in our Jupyter notebook. @@ -335,107 +439,6 @@ df.to_pickle('data/all_years.pkl') ``` -::::::::::::::::::::::::::::::::::::::: challenge - -## Displaying rows and columns - -How would you use slicing and column names to select the following subsets of rows and columns from the circulation DataFrame? - -1. The city column. -2. Rows 10 to 20. -3. Rows 20 to 30 from the zip code column. - -::::::::::::::: solution - -## Solution - -```python -#1 -df['city'] - -#2 -df[10:21] - -#3 -df['zip code'][20:31] - -``` - - -::::::::::::::::::::::::: -:::::::::::::::::::::::::::::::::::::::::::::::::: - - -::::::::::::::::::::::::::::::::::::::: challenge - -## Using loc() - -How would you use `loc()` to select rows 20 to 30 from the zip code column (the same rows as the last example in the challenge above)? - -Tip: slices use "non-inclusive" indexing -- so require you to ask for `df[10:21]` to see row 20, but `loc()` uses inclusive indexing. - -::::::::::::::: solution - -## Solution - -```python -df.loc[20:30, 'zip code'] - -``` -::::::::::::::::::::::::: -:::::::::::::::::::::::::::::::::::::::::::::::::: - -::::::::::::::::::::::::::::::::::::::: challenge - -## Unique items - -How would you display: - -1. all of the unique zip codes in the dataset? -2. the number of unique zip codes in the dataset? - - -::::::::::::::: solution - -## Solution - -```python - -#1 -df['zip code'].unique() - -#2 -df['zip code'].nunique() - -``` -::::::::::::::::::::::::: -:::::::::::::::::::::::::::::::::::::::::::::::::: - -::::::::::::::::::::::::::::::::::::::: challenge - -## Summary statistics and groupby() - -We can apply `mean()` to pandas series' in the same way we used `sum()`, `min()`, and `max()` above. How would you display the following? - -1. the mean number of ytd checkouts grouped by zip code? -2. the mean number of ytd checkouts grouped by zip code, and sorted from smallest to largest? - - -::::::::::::::: solution - -## Solution - -```python -#1 -df.groupby('zip code')['ytd'].mean() - -#2 -df.groupby('zip code')['ytd'].mean().sort_values() - -``` -::::::::::::::::::::::::: -:::::::::::::::::::::::::::::::::::::::::::::::::: - :::::::::::::::::::::::::::::::::::::::: keypoints diff --git a/episodes/tidy.md b/episodes/tidy.md index 993ea974..edef5e0e 100644 --- a/episodes/tidy.md +++ b/episodes/tidy.md @@ -264,6 +264,42 @@ df_long.groupby(['branch', 'month'])['circulation'].agg(['sum', 'mean'])
984 rows × 2 columns
+::::::::::::::::::::::::::::::::::::::: challenge + +## Tidy Data Principles + +How would you reorganize the following table about research data workshops to follow the three tidy data principles? + +1. Every column holds a single variable. +2. Every row represents a single observation. +3. Every cell contains a single value. + +| Date | Length | Content | Instructor | +|------------|---------|-------------|------------| +| 2023-01-15 | 30 min | RDM, DMP | CH | +| 2023-02-02 | 2 hours | Python, RDM | CH, TD | +| 2023-02-03 | 90 min | Python | SP | + +You can use each content unit (e.g., RDM, DMP, Python) as an observation, and breakdown the length of time or instructor initials to match the content unit however you like. + + +::::::::::::::: solution + +## Solution + +| Year | Month | Day | Length (min) | Content | Instructor | +|------|-------|-----|--------------|---------|------------| +| 2023 | 01 | 15 | 20 | RDM | CH | +| 2023 | 01 | 15 | 10 | DMP | CH | +| 2023 | 02 | 02 | 100 | Python | TD | +| 2023 | 02 | 02 | 20 | RDM | CH | +| 2023 | 02 | 03 | 100 | Python | SP | + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + + ## Adding a Date Column In order to plot this data over time in the data visualization we need to do three things to prepare it. First, we need to combine the year and month columns into its own column. Second, convert the new date column to a [datetime](https://docs.python.org/3/library/datetime.html) objec using the Pandas `to_datetime` function. Third, we assign the date column as our index for the data. These steps will set up our data for plotting. @@ -338,41 +374,6 @@ df_long.to_pickle('data/df_long.pkl') ``` ::::::::::::::::::::::::::::::::::::::: challenge -## Tidy Data Principles - -How would you reorganize the following table about research data workshops to follow the three tidy data principles? - -1. Every column holds a single variable. -2. Every row represents a single observation. -3. Every cell contains a single value. - -| Date | Length | Content | Instructor | -|------------|---------|-------------|------------| -| 2023-01-15 | 30 min | RDM, DMP | CH | -| 2023-02-02 | 2 hours | Python, RDM | CH, TD | -| 2023-02-03 | 90 min | Python | SP | - -You can use each content unit (e.g., RDM, DMP, Python) as an observation, and breakdown the length of time or instructor initials to match the content unit however you like. - - -::::::::::::::: solution - -## Solution - -| Year | Month | Day | Length (min) | Content | Instructor | -|------|-------|-----|--------------|---------|------------| -| 2023 | 01 | 15 | 20 | RDM | CH | -| 2023 | 01 | 15 | 10 | DMP | CH | -| 2023 | 02 | 02 | 100 | Python | TD | -| 2023 | 02 | 02 | 20 | RDM | CH | -| 2023 | 02 | 03 | 100 | Python | SP | - -::::::::::::::::::::::::: - -:::::::::::::::::::::::::::::::::::::::::::::::::: - -::::::::::::::::::::::::::::::::::::::: challenge - ## Subsetting df_long Using df_long, create a new DataFrame, `low_circ', that only includes branches with circulation numbers lower than 500 per month. When you create a subset DataFrame, show the following columns: branch, circulation, month, and year. Next, eliminate the rows when the circulation is equal to 0. diff --git a/episodes/variables.md b/episodes/variables.md index 3e4c1d49..c0d77f2b 100644 --- a/episodes/variables.md +++ b/episodes/variables.md @@ -74,6 +74,42 @@ f'{name} is {age} years old' 'Ahmed is 42 years old' ``` +::::::::::::::::::::::::::::::::::::::: challenge + +## F-string Syntax + +Use an f-string to construct output in Python by filling in the blanks with variables and f-string syntax to tell Christina how old she will be in 10 years. + +Tip: You can combine variables and mathematical expressions in an f-string in the same way you can in variable assignment. We'll see more examples of dynamic f-string output as we go through the lesson. + +```python +name = 'Christina' +age = 23 + +f'{____}, you will be ______ in 10 years.' +``` + +::::::::::::::: solution + +## Solution + +```python +f'{name}, you will be {age + 10} in 10 years.' + +``` + +```output +'Christina, you will be 33 in 10 years.' + +``` + + + +::::::::::::::::::::::::: + +:::::::::::::::::::::::::::::::::::::::::::::::::: + + ## Variables must be created before they are used. If a variable doesn't exist yet, or if the name has been misspelled, Python reports an error called a `NameError`. @@ -155,164 +191,116 @@ TypeError Traceback (most recent call last) TypeError: unsupported operand type(s) for -: 'str' and 'str' ``` -## Use an index to get a single character from a string. +::::::::::::::::::::::::::::::::::::::: challenge -We can reference the specific location of a character (individual letters, numbers, and so on) in a string by using its index position. In Python, each character in a string (first, second, etc.) is given a number, which is called an index. Indexes begin from 0 rather than 1. We can use an index in square brackets to refer to the character at that position. +## Fractions -```python -library = 'Alexandria' -library[0] -``` +What type of value is 3.4? +How can you find out? -```output -A -``` +::::::::::::::: solution -## Use a slice to get multiple characters from a string. +## Solution -A slice is a part of a string that we can reference using `[start:stop]`, where `start` is the index of the first character we want and `stop` is the last character. Referencing a string slice does not change the contents of the original string. Instead, the slice returns a copy of the part of the original string we want. +It is a floating-point number (often abbreviated "float"). ```python -library[0:3] +print(type(3.4)) ``` ```output -Ale +