Reinforcement learning 2022 2023: различия между версиями

Материал из Wiki - Факультет компьютерных наук
Перейти к навигации Перейти к поиску
Нет описания правки
Нет описания правки
 
(не показано 11 промежуточных версий этого же участника)
Строка 10: Строка 10:


== About the course ==
== About the course ==
This page contains materials for Mathematical Foundations of Reinforcement learning course in 2021/2022 year, optional one for 2nd year Master students of the Math of Machine Learning program (HSE and Skoltech).
This page contains materials for Mathematical Foundations of Reinforcement learning course in 2022/2023 year, optional one for 2nd year Master students of the Math of Machine Learning program (HSE and Skoltech).


== Grading ==  
== Grading ==  
Строка 17: Строка 17:
* O<sub>Project</sub> for the course project
* O<sub>Project</sub> for the course project
The formula for the final grade is  
The formula for the final grade is  
* O<sub>Final</sub> = 0.5*O<sub>HW</sub> + 0.5*O<sub>Project</sub>
* O<sub>Final</sub> = 0.6*O<sub>HW</sub> + 0.4*O<sub>Project</sub>
with the usual (arithmetical) rounding rule.
with the usual (arithmetical) rounding rule.


[https://docs.google.com/spreadsheets/d/1MPWVIkgxyotHU-P5cE7Gik4C6RTWxTnAVK8Btl7Fw3Y/edit?usp=sharing '''Table with grades''']
[https://docs.google.com/spreadsheets/d/1MPWVIkgxyotHU-P5cE7Gik4C6RTWxTnAVK8Btl7Fw3Y/edit?usp=sharing '''Table with grades''']


== Lectures ==
== Course materials ==
*[https://www.dropbox.com/s/a69ql9duo5jf5gt/Math%20of%20RL%20Lecture%201.pdf?dl=0 ''' Lecture 09.11''']
*[https://www.overleaf.com/read/kbzmvxdzbrxq '''Lectures and seminars notes''']
*[https://www.dropbox.com/s/7zkirk1xykua890/Math_of_RL_Le%20cture_2.pdf?dl=0 ''' Lecture 16.11''']
*[https://colab.research.google.com/drive/10qBq7Ot_1ZpnTeD11P5AnE8jFVj0OLXl?usp=sharing '''Notebook for the first seminar''']
 
== Seminars ==
*[https://www.dropbox.com/s/wc951vseud1q1p2/Seminar_09_11_RL.pdf?dl=0 '''Seminar 09.11'''], [https://www.dropbox.com/s/2h83vbjgew1inen/Seminar_1_RL.mp4?dl=0 '''Seminar 09.11, Video'''], [https://www.dropbox.com/s/bxa8h9vjrnegsql/Bandit_intro_strategies_09_11_2021.ipynb?dl=0 '''Seminar 09.11, Notebook''']
*[https://www.dropbox.com/s/cq0t2o6n4yn6oag/Seminar_16_11_RL.mp4?dl=0 '''Seminar 16.11, Video'''],
*[https://www.dropbox.com/s/ex8v9w3smar70m7/Seminar_23_11_RL.mp4?dl=0 '''Seminar 23.11, Video'''],
*[https://www.dropbox.com/s/v1ywnk8eyhourjq/Seminar_07_12_RL.mp4?dl=0 '''Seminar 07.12, Video'''],


== Recommended literature ==
== Recommended literature ==
'''Lecture and seminar 09.11'''


* Sebastien Bubek, Nicolo Cesa-Bianchi. Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems. Chapter 2. http://sbubeck.com/SurveyBCB12.pdf
* Sebastien Bubek, Nicolo Cesa-Bianchi. Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems. Chapter 2. http://sbubeck.com/SurveyBCB12.pdf
* Richard S. Sutton, Andrew G. Barto. Reinforcement Learning: An Introduction. Chapter 2. http://incompleteideas.net/book/the-book-2nd.html;
* Richard S. Sutton, Andrew G. Barto. Reinforcement Learning: An Introduction. Chapter 2. http://incompleteideas.net/book/the-book-2nd.html;
* Botao Hao et al. Bootstrapping Upper Confidence Bound. https://arxiv.org/abs/1906.05247
* Botao Hao et al. Bootstrapping Upper Confidence Bound. https://arxiv.org/abs/1906.05247
 
* Aleksandrs Slivkins. Introduction to Multi-Armed Bandits. https://arxiv.org/abs/1904.07272 [Chapter 1]
'''Lecture and seminar 16.11'''
*[https://www.dropbox.com/s/wc951vseud1q1p2/Seminar_09_11_RL.pdf?dl=0 '''Seminar 09.11'''], [https://www.dropbox.com/s/2h83vbjgew1inen/Seminar_1_RL.mp4?dl=0 '''Seminar 09.11, Video'''],


==Homeworks ==
==Homeworks ==
*[https://www.dropbox.com/s/k2at9lixvshpcbw/HW_1_RL_2021.pdf?dl=0 '''Homework №1, deadline 19.12.2021, 23:59'''], [https://www.dropbox.com/s/l7pma6kwnopl856/HW_1_task_2.ipynb?dl=0 '''Environment for task №2'''],
*[https://github.com/svsamsonov/Math_RL_2022_2023 '''HW #1, deadline: 04.12.22, 23:59''']
*[https://www.dropbox.com/s/jynwji3dw3xxjww/HW_2_RL_2021.pdf?dl=0 '''Homework №2, deadline 19.12.2021, 23:59'''].


== Projects ==
== Projects ==

Текущая версия от 11:39, 21 ноября 2022

Lecturers and Seminarists

Lecturer Alexey Naumov [anaumov@hse.ru] T924
Seminarist Sergey Samsonov [svsamsonov@hse.ru] T926

About the course

This page contains materials for Mathematical Foundations of Reinforcement learning course in 2022/2023 year, optional one for 2nd year Master students of the Math of Machine Learning program (HSE and Skoltech).

Grading

The final grade consists of 2 components (each is non-negative real number from 0 to 10, without any intermediate rounding) :

  • OHW for the hometasks
  • OProject for the course project

The formula for the final grade is

  • OFinal = 0.6*OHW + 0.4*OProject

with the usual (arithmetical) rounding rule.

Table with grades

Course materials

Recommended literature

Homeworks

Projects