Data Science Weekly - Data Science Weekly - Issue 475

Curated news, articles and jobs related to Data Science.
Keep up with all the latest developments

Email not displaying correctly?
View it in your browser.

Issue #475

December 29 2022

Editor's Picks

Jürgen Schmidhuber's Annotated History of Modern AI and Deep Learning
Machine learning is the science of credit assignment: finding patterns in observations that predict the consequences of actions and help to improve future performance. Credit assignment is also required for human understanding of how the world works, not only for individuals navigating daily life, but also for academic professionals like historians who interpret the present in light of past events. Here I focus on the history of modern artificial intelligence (AI) which is dominated by artificial neural networks (NNs) and deep learning, both conceptually closer to the old field of cybernetics than to what's been called AI since 1956...

Cats, Pi, and Machine Learning [Project]
There is a cat that wanders around my neighbourhood. I wanted to build something that would notify me whenever it came to my backyard...I thought to myself, if I can get a picture of my backyard as the input I can process the image by checking if there is a cat in it and send a notification to my phone as the output...For the picture input, I attached a camera module to a raspberry pi 4. For the processing, I wrote some code that would run on the pi, which would periodically take an image and run object detection on it. For the output, I relied on sending a message via Signal to my phone...

Another Year of #TidyTuesday
After another 52 data visualisations created for #TidyTuesday, it's time for the annual round-up! Read this blog post for some interesting R packages discovered, a few new I've tricks learnt, and the data visualisations I'd like to do again...

A Message from this week's Sponsor:

Ilum the Spark cluster manager and monitoring tool

With Ilum's solution, everyone can now quickly and easily deploy Apache Spark on any Kubernetes cluster. Our software eliminates the need for tedious configuration and reduces the time needed for deployment from days to minutes. By leveraging the power of container orchestration and Apache Spark's scalability and reliability, we are making it easier than ever to stay ahead of the curve and explore the future of Big Data.

Ilum provides an all-in-one solution for:

Data Science on Kubernetes
Hadoop replacement
Apache Livy alternative
Integration with Jupyter and Apache Zeppelin

It's free! Unlock the power of Big Data today with Ilum.

Learn more about Ilum!

Data Science Articles & Videos

Annotated Forest Plots using ggplot2
rstats blog post! a walkthrough on making forest plots and adding effect size annotations using ggplot2...

The Build vs. Buy Guide for the Modern Data Stack
Nishith Agarwal, Head of Data & ML Platforms at Lyra Health and creator of the popular open source data management framework, Apache Hudi, outlines his blueprint for building (and buying) the data stack of your dreams. In Part I of this series, Nishith discusses some initial considerations and shares a framework for getting started...

Forward-Forward Algorithm App
This app implements a complete open-source version of Geoffrey Hinton's Forward Forward Algorithm, an alternative approach to backpropagation...The Forward Forward algorithm is a method for training deep neural networks that replaces the backpropagation forward and backward passes with two forward passes, one with positive (i.e., real) data and the other with negative data that could be generated by the network itself...

Large Language Models Encode Clinical Knowledge
We present MultiMedQA, a benchmark combining six existing open question answering datasets spanning professional medical exams, research, and consumer queries; and HealthSearchQA, a new free-response dataset of medical questions searched online...we evaluate PaLM (a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM, on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA, MedMCQA, PubMedQA, MMLU clinical topics), including 67.6% accuracy on MedQA (US Medical License Exam questions), surpassing prior state-of-the-art by over 17%...

Partially renaming columns using a lookup table
Usually data sets come with short column names, which makes it easy to clean and manipulate the data. However, when presenting the data to stakeholders, in form of tables or plots, we often need longer, meaningful names. In many cases we have a lookup table which contains long and short versions of the column names so that we can “easily” replace the names when needed...Below we’ll look at how to rename columns using different approaches in R...

ChatBCG: Generative AI for Slides [Twitter Thread]
This Christmas Joseph Semrai and I finally got it working!! After DALL-E 2 for images and ChatGPT for text, the final step to make all of us redundant: The world’s first Text-to-PowerPoint AI...

Top Python libraries of 2022
Welcome to the 8th edition of our Top Python Libraries list!...We are excited to present this year's picks for the most innovative developments in the Python ecosystem. From this edition, we are expanding our list to include not only libraries per-se, but also tools that are built to belong in the Python ecosystem — some of which are not written in Python as you’ll see...

Ask HN: Upskilling as a Data Engineer [HN Discussion]
What should a data engineer learn as a part of Upskilling in 2022/2023?...New languages like Rust/Ocaml/Nim...if yes then which?...I don't think learning an ETL tool will be helpful because essentially they are all one and the same...Any tips?...

Vanishing Gradients Podcast #15: Uncertainty, Risk, and Simulation in Data Science
Hugo speaks with JD Long, agricultural economist, quant, and stochastic modeler, about decision making under uncertainty and how we can use our knowledge of risk, uncertainty, probabilistic thinking, causal inference, and more to help us use data science and machine learning to make better decisions in an uncertain world...This is part 1 of a two part conversation. In this, part 1, we discuss risk, uncertainty, probabilistic thinking, and simulation, all with a view towards improving decision making and we draw on examples from our personal lives, the pandemic, our jobs, the reinsurance space, and the corporate world. In part 2, we’ll get into the nitty gritty of decision making under uncertainty...

University Professor Catches Student Cheating With ChatGPT
A South Carolina college professor is raising the alarm after discovering that ChatGPT, a chatbot with artificial intelligence, was being used by a student to create an essay....

At the forefront of AI with Hugo Larochelle, Research Scientist at Google
On the show, we spoke about: His love of teaching and starting a YouTube channel. The rapid advances of the past 12 months. The illusion of AGI Enabling new scientific discoveries, Philanthropy, and the generous $1M donation from Hugo and his wife to combat climate change...

What tech books (or courses?) have really helped you this past year? [Twitter Thread]
It's the end of the year and some people have education budgets to spend! What tech books (or courses?) have really helped you this past year?...

Tool*

Build powerful ML visualizations with Comet

With just 2 lines of code, Comet automatically logs metrics, hyperparameters, libraries, and more. This means automatic chart generation so you can easily manage training runs in real time. When you combine that with:

built-in visualizations (like the image panel),
custom project views, and
your own python panels,

Comet is a powerful tool for optimizing your ML workflow. All for free! Less friction, more ML.

Create your free account.

*Sponsored post. If you want to be featured here, or as our main sponsor, contact us!

Tool*

Now where can I find that query... 🔍🔍

Did you put it in a doc? Slack? Teams? Notes?? Make searching for a query a thing of the past with Sherloq. Sherloq helps data analysts save, organize, and share their metrics, most used or complex queries for seamless collaboration within their organization. It’s a secure add-on (no integrations or permissions necessary) that works on the popular query editors. Start organizing your query repository in a shared workspace with Sherloq beta.

Try Sherloq For Free

*Sponsored post. If you want to be featured here, or as our main sponsor, contact us!

Jobs

Data Scientist / Machine Learning Engineer - Epsilon - NYC

Epsilon Strategy and Insights, Data Sciences team is looking for a talented team player in a Data Scientist/Machine Learning Engineer role. You are an expert, mentor and advocate. You have strong machine learning and deep learning background and are passionate about transforming data into ml models. You welcome the challenge of data science and are proficient in Python, Spark MLLib, Tensorflow, Keras, ML algorithms and Deep Neural Networks, Big Data. You must be self-driven, take initiative and want to work in a dynamic, busy and innovative group...

Want to post a job here? Email us for details --> team@datascienceweekly.org

Training & Resources

Transformers from Scratch
I procrastinated a deep dive into transformers for a few years. Finally the discomfort of not knowing what makes them tick grew too great for me. Here is that dive...Transformers were introduced in this 2017 paper as a tool for sequence transduction—converting one sequence of symbols to another. The most popular examples of this are translation, as in English to German. It has also been modified to perform sequence completion—given a starting prompt, carry on in the same vein and style. They have quickly become an indispensible tool for research and product development in natural language processing...

Xavier Glorot Initialization in Neural Networks — Math Proof
Detailed derivation for finding optimal initial distributions of weight matrices in deep learning layers with tanh activation function...

An overview of gradient descent optimization algorithms
Gradient descent is the preferred way to optimize neural networks and many other machine learning algorithms but is often used as a black box. This post explores how many of the most popular gradient-based optimization algorithms such as Momentum, Adagrad, and Adam actually work...

Last Week's Newsletter's 3 Most Clicked Links

Why Business Data Science Irritates Me

Everything I learned about accidentally running a successful tech conference

Build a GPT-3 app: How I used GPT-3 to make gifting easier

* Based on unique clicks.
** Find last week's newsletter here.

Cutting Room Floor

P.S., Enjoy the newsletter? Please forward it to your friends and colleagues - we'd love to have them onboard :) All the best, Hannah & Sebastian

Follow on Twitter

unsubscribe from this list update subscription preferences

Data Science Weekly - Data Science Weekly - Issue 475

Issue #475

December 29 2022

Editor's Picks

A Message from this week's Sponsor:

Data Science Articles & Videos

Tool*

Tool*

Jobs

Training & Resources

Last Week's Newsletter's 3 Most Clicked Links

Cutting Room Floor

Older messages

Data Science Weekly - Issue 474

Data Science Weekly - Issue 473

Data Science Weekly - Issue 472

Data Science Weekly - Issue 471

Data Science Weekly - Issue 470

You Might Also Like

Import AI 399: 1,000 samples to make a reasoning model; DeepSeek proliferation; Apple's self-driving car simulator

Defining Your Paranoia Level: Navigating Change Without the Overkill

5 ways AI can help with taxes 🪄

Recurring Automations + Secret Updates

The First Provable AI-Proof Game: Introducing Butterfly Wings 4

GCP Newsletter #437

Charted | The 1%'s Share of U.S. Wealth Over Time (1989-2024) 💰

The Great Social Media Diaspora & Tapestry is here

Daily Coding Problem: Problem #1689 [Medium]

📧 Stop Conflating CQRS and MediatR

Data Science Weekly - Data Science Weekly - Issue 475

Issue #475 December 29 2022

Editor's Picks

A Message from this week's Sponsor:

Data Science Articles & Videos

Tool*

Tool*

Jobs

Training & Resources

Last Week's Newsletter's 3 Most Clicked Links

Cutting Room Floor

Older messages

You Might Also Like

Issue #475

December 29 2022