[SPARK-4583] [mllib] LogLoss for GradientBoostedTrees fix + doc updates #3439

jkbradley · 2014-11-25T02:14:51Z

Currently, the LogLoss used by GradientBoostedTrees has 2 issues:

the gradient (and therefore loss) does not match that used by Friedman (1999)
the error computation uses 0/1 accuracy, not log loss

This PR updates LogLoss.
It also adds some doc for boosting and forests.

I tested it on sample data and made sure the log loss is monotonically decreasing with each boosting iteration.

CC: @mengxr @manishamde @codedeft

SparkQA · 2014-11-25T02:20:35Z

Test build #23811 has started for PR 3439 at commit 7c38962.

This patch merges cleanly.

SparkQA · 2014-11-25T03:42:47Z

Test build #23811 has finished for PR 3439 at commit 7c38962.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

AmplabJenkins · 2014-11-25T03:42:50Z

Test PASSed.
Refer to this link for build results (access rights to CI server needed):
https://amplab.cs.berkeley.edu/jenkins//job/SparkPullRequestBuilder/23811/
Test PASSed.

mengxr · 2014-11-25T10:10:42Z

mllib/src/main/scala/org/apache/spark/mllib/tree/loss/SquaredError.scala

MSE is not usually defined with multiplier 1/2. Shall we use a different name here, or example, mean squared loss or average loss?

I'll remove the 1/2. It's probably better to have an odd loss (which only experts need to know about) than to have an odd name (which everyone needs to recognize).

…orests, and boosting

…test suite since it effectively doubles the gradient and loss. * Added doc for developers within RandomForest. * Small cleanup in test suite (generating data only once)

jkbradley · 2014-11-25T21:31:04Z

I just pushed an update which includes:

removing the 1/2 from SquaredError. This also required updating the test suite since it effectively doubles the gradient and loss.
Added doc for developers within RandomForest.
small cleanup in test suite (generating data only once)

SparkQA · 2014-11-25T21:32:51Z

Test build #23849 has started for PR 3439 at commit 5e52bff.

This patch merges cleanly.

SparkQA · 2014-11-25T22:58:16Z

Test build #23849 has finished for PR 3439 at commit 5e52bff.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

AmplabJenkins · 2014-11-25T22:58:20Z

Test PASSed.
Refer to this link for build results (access rights to CI server needed):
https://amplab.cs.berkeley.edu/jenkins//job/SparkPullRequestBuilder/23849/
Test PASSed.

manishamde · 2014-11-26T00:37:02Z

@jkbradley I am trying to find my reference for the LogLoss calculations.

mengxr · 2014-11-26T00:46:16Z

mllib/src/main/scala/org/apache/spark/mllib/tree/loss/LogLoss.scala

There is an issue with numerical stability. Maybe we can fix it in this PR. The problem appears when w = -2.0 * point.label * prediction is large. math.exp(w) would overflow while math.log(1 + math.exp(w)) should be close to w. When w < 0, we can use

math.log1p(math.exp(w)).

Otherwise, we should use

w + math.log1p(math.exp(-w))

manishamde · 2014-11-26T00:56:45Z

@jkbradley LGTM. Thanks for the documentation too -- it is really helpful.

SparkQA · 2014-11-26T01:30:40Z

Test build #23854 has started for PR 3439 at commit ed5da2c.

This patch merges cleanly.

jkbradley · 2014-11-26T01:35:25Z

Updated LogLoss.
@mengxr @manishamde Thanks for looking at this!

SparkQA · 2014-11-26T01:40:19Z

Test build #23856 has started for PR 3439 at commit a27eb6d.

This patch merges cleanly.

SparkQA · 2014-11-26T01:41:27Z

Test build #23856 has finished for PR 3439 at commit a27eb6d.

This patch fails Scala style tests.
This patch merges cleanly.
This patch adds no public classes.

AmplabJenkins · 2014-11-26T01:41:28Z

Test FAILed.
Refer to this link for build results (access rights to CI server needed):
https://amplab.cs.berkeley.edu/jenkins//job/SparkPullRequestBuilder/23856/
Test FAILed.

SparkQA · 2014-11-26T02:07:36Z

Test build #23862 has started for PR 3439 at commit cfec17e.

This patch merges cleanly.

SparkQA · 2014-11-26T02:56:11Z

Test build #23854 has finished for PR 3439 at commit ed5da2c.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

AmplabJenkins · 2014-11-26T02:56:14Z

Test PASSed.
Refer to this link for build results (access rights to CI server needed):
https://amplab.cs.berkeley.edu/jenkins//job/SparkPullRequestBuilder/23854/
Test PASSed.

SparkQA · 2014-11-26T03:40:46Z

Test build #23862 has finished for PR 3439 at commit cfec17e.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

AmplabJenkins · 2014-11-26T03:40:49Z

Test PASSed.
Refer to this link for build results (access rights to CI server needed):
https://amplab.cs.berkeley.edu/jenkins//job/SparkPullRequestBuilder/23862/
Test PASSed.

mengxr · 2014-11-26T04:10:53Z

LGTM. Merged into master and branch-1.2. Thanks!

Currently, the LogLoss used by GradientBoostedTrees has 2 issues: * the gradient (and therefore loss) does not match that used by Friedman (1999) * the error computation uses 0/1 accuracy, not log loss This PR updates LogLoss. It also adds some doc for boosting and forests. I tested it on sample data and made sure the log loss is monotonically decreasing with each boosting iteration. CC: mengxr manishamde codedeft Author: Joseph K. Bradley <[email protected]> Closes apache#3439 from jkbradley/gbt-loss-fix and squashes the following commits: cfec17e [Joseph K. Bradley] removed forgotten temp comments a27eb6d [Joseph K. Bradley] corrections to last log loss commit ed5da2c [Joseph K. Bradley] updated LogLoss (boosting) for numerical stability 5e52bff [Joseph K. Bradley] * Removed the 1/2 from SquaredError. This also required updating the test suite since it effectively doubles the gradient and loss. * Added doc for developers within RandomForest. * Small cleanup in test suite (generating data only once) e57897a [Joseph K. Bradley] Fixed LogLoss for GradientBoostedTrees, and updated doc for losses, forests, and boosting (cherry picked from commit c251fd7) Signed-off-by: Xiangrui Meng <[email protected]>

…+ doc updates We reverted #3439 in branch-1.2 due to missing `import o.a.s.SparkContext._`, which is no longer needed in master (#3262). This PR adds #3439 back to branch-1.2 with correct imports. Github is out-of-sync now. The real changes are the last two commits. Author: Joseph K. Bradley <[email protected]> Author: Xiangrui Meng <[email protected]> Closes #3474 from mengxr/SPARK-4583-1.2 and squashes the following commits: aca2abb [Xiangrui Meng] add import o.a.s.SparkContext._ for v1.2 6b5564a [Joseph K. Bradley] [SPARK-4583] [mllib] LogLoss for GradientBoostedTrees fix + doc updates

mengxr reviewed Nov 25, 2014
View reviewed changes

jkbradley added 2 commits November 25, 2014 12:37

Fixed LogLoss for GradientBoostedTrees, and updated doc for losses, f…

e57897a

…orests, and boosting

* Removed the 1/2 from SquaredError. This also required updating the …

5e52bff

…test suite since it effectively doubles the gradient and loss. * Added doc for developers within RandomForest. * Small cleanup in test suite (generating data only once)

jkbradley force-pushed the gbt-loss-fix branch from 7c38962 to 5e52bff Compare November 25, 2014 21:30

mengxr reviewed Nov 26, 2014
View reviewed changes

updated LogLoss (boosting) for numerical stability

ed5da2c

corrections to last log loss commit

a27eb6d

removed forgotten temp comments

cfec17e

mengxr mentioned this pull request Nov 26, 2014

[BRANCH-1.2][SPARK-4583][MLLIB] LogLoss for GradientBoostedTrees fix + doc updates #3474

Closed

asfgit closed this in c251fd7 Nov 26, 2014

jkbradley deleted the gbt-loss-fix branch December 4, 2014 20:27

[SPARK-4583] [mllib] LogLoss for GradientBoostedTrees fix + doc updates #3439

[SPARK-4583] [mllib] LogLoss for GradientBoostedTrees fix + doc updates #3439

Uh oh!

Conversation

jkbradley commented Nov 25, 2014

Uh oh!

SparkQA commented Nov 25, 2014

Uh oh!

SparkQA commented Nov 25, 2014

Uh oh!

AmplabJenkins commented Nov 25, 2014

Uh oh!

mengxr Nov 25, 2014

Choose a reason for hiding this comment

Uh oh!

jkbradley Nov 25, 2014

Choose a reason for hiding this comment

Uh oh!

jkbradley commented Nov 25, 2014

Uh oh!

SparkQA commented Nov 25, 2014

Uh oh!

SparkQA commented Nov 25, 2014

Uh oh!

AmplabJenkins commented Nov 25, 2014

Uh oh!

manishamde commented Nov 26, 2014

Uh oh!

mengxr Nov 26, 2014

Choose a reason for hiding this comment

Uh oh!

jkbradley Nov 26, 2014

Choose a reason for hiding this comment

Uh oh!

manishamde commented Nov 26, 2014

Uh oh!

SparkQA commented Nov 26, 2014

Uh oh!

jkbradley commented Nov 26, 2014

Uh oh!

SparkQA commented Nov 26, 2014

Uh oh!

SparkQA commented Nov 26, 2014

Uh oh!

AmplabJenkins commented Nov 26, 2014

Uh oh!

SparkQA commented Nov 26, 2014

Uh oh!

SparkQA commented Nov 26, 2014

Uh oh!

AmplabJenkins commented Nov 26, 2014

Uh oh!

SparkQA commented Nov 26, 2014

Uh oh!

AmplabJenkins commented Nov 26, 2014

Uh oh!

mengxr commented Nov 26, 2014

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants