久久综合色88_欧美激情国产日韩精品一区18_午夜精品一区二区三区在线观看 _自拍日韩亚洲一区在线

課程目錄: 基于函數(shù)逼近的預(yù)測與控制培訓(xùn)
4401 人關(guān)注
(78637/99817)
課程大綱:

    基于函數(shù)逼近的預(yù)測與控制培訓(xùn)

 

 

 

Welcome to the Course!

Welcome to the third course in the Reinforcement Learning Specialization:

Prediction and Control with Function Approximation, brought to you by the University of Alberta,

Onlea, and Coursera.

In this pre-course module, you'll be introduced to your instructors,

and get a flavour of what the course has in store for you.

Make sure to introduce yourself to your classmates in the "Meet and Greet" section!

On-policy Prediction with Approximation

This week you will learn how to estimate a value function for a given policy,

when the number of states is much larger than the memory available to the agent.

You will learn how to specify a parametric form of the value function,

how to specify an objective function, and how estimating gradient descent can be used to estimate values from interaction with the world.

Constructing Features for Prediction

The features used to construct the agent’s value estimates are perhaps the most crucial part of a successful learning system.

In this module we discuss two basic strategies for constructing features: (1) fixed basis that form an exhaustive partition of the input,

and (2) adapting the features while the agent interacts with the world via Neural Networks and Backpropagation.

In this week’s graded assessment you will solve a simple but infinite state prediction task with a Neural Network and

TD learning.Control with ApproximationThis week,

you will see that the concepts and tools introduced in modules two and three allow straightforward extension of classic

TD control methods to the function approximation setting. In particular,

you will learn how to find the optimal policy in infinite-state MDPs by simply combining semi-gradient

TD methods with generalized policy iteration, yielding classic control methods like Q-learning, and Sarsa.

We conclude with a discussion of a new problem formulation for RL---average reward---which will undoubtedly

be used in many applications of RL in the future.

Policy GradientEvery algorithm you have learned about so far estimates

a value function as an intermediate step towards the goal of finding an optimal policy.

An alternative strategy is to directly learn the parameters of the policy.

This week you will learn about these policy gradient methods, and their advantages over value-function based methods.

You will also learn how policy gradient methods can be used

to find the optimal policy in tasks with both continuous state and action spaces.

主站蜘蛛池模板: 热久久免费国产视频| 欧美一区二区三区精品电影| 中文字幕99| 激情五月六月婷婷| 狠狠色狠狠色综合人人| 蜜桃av噜噜一区二区三区| 日本精品免费视频 | 日韩不卡av| 日本精品一区在线观看| 欧美国产日韩在线播放| 国产专区精品视频| 91禁国产网站| 日韩国产欧美亚洲| 久久精品国产精品亚洲| 国产精品网站免费| 伊人久久大香线蕉成人综合网| 欧美日本韩国国产| 国产精品 欧美在线| 色乱码一区二区三在线看| 久章草在线视频| 国产mv久久久| 日韩视频精品在线| 国产欧美日韩中文字幕在线| 国产欧美欧洲| 久久久久久九九| 91精品视频在线免费观看| 久久人人爽亚洲精品天堂| 国产精品久久久久久av| 欧美日韩一区在线视频| 国产专区欧美专区| 91精品国产91久久久久福利| 久久免费福利视频| 国产一级片91| 尤物一区二区三区| 人妻少妇精品无码专区二区| 奇米影视首页 狠狠色丁香婷婷久久综合| 在线视频一区观看| 欧美精品性视频| 国产精品久久久999| 青春草国产视频| 国产精品久久久久久亚洲影视 |