mindspore.experimental.optim.Adagrad

class mindspore.experimental.optim.Adagrad(params, lr=1e-2, lr_decay=0.0, weight_decay=0.0, initial_accumulator_value=0.0, eps=1e-10, *, maximize=False)[source]

Implements Adagrad algorithm.

\begin{array}{r} \begin{aligned} input : γ (lr), θ_{0} (params), f (θ) (objective), λ (weight decay), \\ τ (initial accumulator value), η (lr decay) \\ initialize : s t a t e_s u m_{0} \leftarrow 0 \\ for t = 1 to \dots do \\ g_{t} \leftarrow \nabla_{θ} f_{t} (θ_{t - 1}) \\ \tilde{γ} \leftarrow γ / (1 + (t - 1) η) \\ if λ \neq 0 \\ g_{t} \leftarrow g_{t} + λ θ_{t - 1} \\ s t a t e_s u m_{t} \leftarrow s t a t e_s u m_{t - 1} + g_{t}^{2} \\ θ_{t} \leftarrow θ_{t - 1} - \tilde{γ} \frac{g_{t}}{\sqrt{s t a t e_s u m_{t}} + ϵ} \\ r e t u r n θ_{t} \end{aligned} \end{array}

For more details about Adagrad algorithm, please refer to Adaptive Subgradient Methods for Online Learning and Stochastic Optimization.

Warning

This is an experimental optimizer API that is subject to change. This module must be used with lr scheduler module in LRScheduler Class .

Parameters

params (Union[list(Parameter), list(dict)]) – list of parameters to optimize or dicts defining parameter groups.
lr (Union[int, float, Tensor], optional) – learning rate. Default: 1e-2.
lr_decay (float, optional) – learning rate decay. Default: 0..
weight_decay (float, optional) – weight decay (L2 penalty). Default: 0..
initial_accumulator_value (float, optional) – the initial accumulator value. Default: 0..
eps (float, optional) – term added to the denominator to improve numerical stability. Default: 1e-10.

Keyword Arguments

maximize (bool, optional) – maximize the params based on the objective, instead of minimizing. Default: False.

Inputs:

gradients (tuple[Tensor]) - The gradients of params.

Raises

ValueError – If the learning rate is not int, float or Tensor.
ValueError – If the learning rate is less than 0.
ValueError – If the learning rate decay is less than 0.
ValueError – If the weight_decay is less than 0.
ValueError – If the initial_accumulator_value is less than 0.0.
ValueError – If the eps is less than 0.0.

Supported Platforms:: Ascend GPU CPU

Examples

>>> import mindspore
>>> from mindspore import nn
>>> from mindspore.experimental import optim
>>> # Define the network structure of LeNet5. Refer to
>>> # https://gitee.com/mindspore/docs/blob/master/docs/mindspore/code/lenet.py
>>> net = LeNet5()
>>> loss_fn = nn.SoftmaxCrossEntropyWithLogits(sparse=True)
>>> optimizer = optim.Adagrad(net.trainable_params(), lr=0.1)
>>> def forward_fn(data, label):
...     logits = net(data)
...     loss = loss_fn(logits, label)
...     return loss, logits
>>> grad_fn = mindspore.value_and_grad(forward_fn, None, optimizer.parameters, has_aux=True)
>>> def train_step(data, label):
...     (loss, _), grads = grad_fn(data, label)
...     optimizer(grads)
...     return loss