神经网络中的矩阵求导及反向传播推导

阅读量：321 次

发布时间：2019-03-03

本文共 1475 字，大约阅读时间需要 4 分钟。

两层全连接神经网络的实现, 包括网络的实现、梯度的反向传播计算和权重更新过程：

# -*- coding: utf-8 -*-import numpy as np# N is batch size; D_in is input dimension;# H is hidden dimension; D_out is output dimension.N, D_in, H, D_out = 64, 1000, 100, 10# Create random input and output datax = np.random.randn(N, D_in)y = np.random.randn(N, D_out)# Randomly initialize weightsw1 = np.random.randn(D_in, H)w2 = np.random.randn(H, D_out)learning_rate = 1e-6for t in range(500):    # Forward pass: compute predicted y    h = x.dot(w1)    h_relu = np.maximum(h, 0)    y_pred = h_relu.dot(w2)    # Compute and print loss    loss = np.square(y_pred - y).sum()    print(t, loss)    # Backprop to compute gradients of w1 and w2 with respect to loss    grad_y_pred = 2.0 * (y_pred - y)    grad_w2 = h_relu.T.dot(grad_y_pred)    grad_h_relu = grad_y_pred.dot(w2.T)    grad_h = grad_h_relu.copy()    grad_h[h < 0] = 0    grad_w1 = x.T.dot(grad_h)    # Update weights    w1 -= learning_rate * grad_w1    w2 -= learning_rate * grad_w2

这里解决了我一个错误的认知：以为最速下降法跟各个变量计算的导数无关，而其实就是每个变量各自按自己的导数下降就可以实现函数最陡的坡进行下降；在图形上可以理解多个向量合并成一个方向；

反向传播过程，核心代码如下

h = x.dot(w1)h_relu = np.maximum(h, 0)y_pred = h_relu.dot(w2)loss = np.square(y_pred - y).sum()grad_y_pred = 2.0 * (y_pred - y)    # 64 x 10grad_w2 = h_relu.T.dot(grad_y_pred) # 100 x 10grad_h_relu = grad_y_pred.dot(w2.T) # 64 x 100grad_h = grad_h_relu.copy()         # 64 x 100grad_h[h < 0] = 0                   # 64 x 100grad_w1 = x.T.dot(grad_h)           # 1000 x 100

问题：如何实现relu求导呢？

转载地址：http://usgm.baihongyu.com/

你可能感兴趣的文章

MySQL Cluster 7.0.36 发布

Multimodal Unsupervised Image-to-Image Translation多通道无监督图像翻译

multipart/form-data与application/octet-stream的区别、application/x-www-form-urlencoded

mysql cmake 报错,MySQL云服务器应用及cmake报错解决办法

Multiple websites on single instance of IIS

mysql CONCAT()函数拼接有NULL

multiprocessing.Manager 嵌套共享对象不适用于队列

multiprocessing.pool.map 和带有两个参数的函数

MYSQL CONCAT函数

multiprocessing.Pool:map_async 和 imap 有什么区别?

MySQL Connector/Net 句柄泄露

multiprocessor（中)

mysql CPU使用率过高的一次处理经历

Multisim中555定时器使用技巧

MySQL CRUD 数据表基础操作实战

multisim变压器反馈式_穿过隔离栅供电：认识隔离式直流/ 直流偏置电源

mysql csv import meets charset

multivariate_normal TypeError: ufunc ‘add‘ output (typecode ‘O‘) could not be coerced to provided……

MySQL DBA 数据库优化策略

multi_index_container