顯示具有 Machine Learning 標籤的文章。 顯示所有文章
顯示具有 Machine Learning 標籤的文章。 顯示所有文章

2016年7月10日 星期日

[ML] logistic gradient accent and Andrew Ng gradient decent


ref:

http://www.csdn123.com/html/topnews201408/66/10366.htm


the gradient fun is l(n)

normally hope probability biggest, so use gradient Accent to find largest theta of l(n)

but

Andrew Ng set J(theta)=-(1/m)l(n), there has a "-", so become find the minimum , so
use gradient decent.


2016年5月25日 星期三

[Machine Learning] 'Training/Cross-Validation/Test

from:http://stats.stackexchange.com/questions/19048/what-is-the-difference-between-test-set-and-validation-set

The concept of 'Training/Cross-Validation/Test' Data Sets is as simple as this. When you have a large data set, it's recommended to split it into 3 parts:

++Training set (60% of the original data set): This is used to build up our prediction algorithm. Our algorithm tries to tune itself to the quirks of the training data sets. In this phase we usually create multiple algorithms in order to compare their performances during the Cross-Validation Phase.

++Cross-Validation set (20% of the original data set): This data set is used to compare the performances of the prediction algorithms that were created based on the training set. We choose the algorithm that has the best performance.

++Test set (20% of the original data set): Now we have chosen our preferred prediction algorithm but we don't know yet how it's going to perform on completely unseen real-world data. So, we apply our chosen prediction algorithm on our test set in order to see how it's going to perform so we can have an idea about our algorithm's performance on unseen data.

Notes:

-It's very important to keep in mind that skipping the test phase is not recommended, because the algorithm that performed well during the cross-validation phase doesn't really mean that it's truly the best one, because the algorithms are compared based on the cross-validation set and its quirks and noises...

-During the Test Phase, the purpose is to see how our final model is going to deal in the wild, so in case its performance is very poor we should repeat the whole process starting from the Training Phase.

2016年5月24日 星期二

[Machine Learing] Training set 6: cross validation set 2: testing set 2

我们首先定义train、cv(cross validation)和test error。
之后我们根据train error训练模型。如果是模型比较时,我们就训练出多个不同的模型,如上图。
之后,对于模型选择。我们让cross validation set通过的这些模型,并计算这些模型的error(cv error)。我们选cv error最小的那个模型作为最好的模型。
最后,对于估计模型的performance(也就是generalization):我们让   test set通过我们刚刚用cv error选出的模型,得到test error,作为performance。
需要强调的是,如果不做模型选择,那么就不需要validation set了。只要training和test set。两者的比例大概是7:3。
然后,这里解释下为什么做模型选择要validation set和test set分开。首先我们不能拿training set做模型选择,因为那样没有generalization了(我们的算法就是根据training set来优化模型的,所以没有再用这个set做模型选择没意义)。其次,如果我们用test set选模型,那么选择的模型就fit到 test set,结果我们同样失去了generalization。所以要用validation set选模型。用test set测generalization。
=>
Use cross validation set to chose polynomial model (theta from training set), use testing set data to judge the performance of model which you just get. 

****
If you do not have to select model, then just use training set 7 : testing set 3 .

2016年5月3日 星期二

[octave] build octave under sublime text 3 (pause not work)

tools->build system->new build system

    {
        "path": "/usr/bin/",//path of your Octave bin folder
        "cmd": ["octave", "--no-site-file", "-p $file_path", "$file_name"]
    }

2016年4月28日 星期四

[Machine Learning][Study Note][Week1] Matrix


可以利用矩陣運算, 將同一組資料對於不同多個h(x)一次帶入,非常方便

 


一些其他特性


矩陣 I 如果為Identity, A*I=I*A, 則有交換律
否則一般矩陣 A*B != B*A

矩陣有結合律
A+(B+C)=(A+B)+C


可以用軟體求Inverse

A*A(inverse)=I

ex: 3*1/3=1, 則1/3就是 3 的inverse.  note: 0 沒有 inverse.

















[Machine Learning][Study Note][Week1] Cost Function & Gradient Descent

一個線性回歸的公式

h(x)=θ0+θ1*x

h(x)=hypothesis



 求 Cost Function 的最小值, 也就是說, 要對 h(x) 找cost 最小的,此h(x) 就會當作

我們的回歸公式

所以, 我們要決定怎樣的東西來當作cost function? 這裡我們用距離的平方,
所以, 找一個一階線性回歸的最佳解, 就是這條線能讓所有訓練資料帶進去
得到所有人距離這條線的總和最小, 那就把係數θ0, θ1拿來用.




Gradient Descent
=你把普通的兩變數函數畫出來後, X=θ0,Y=θ1,Z=Cost,
你就想像你起始站在山頂上, 但這坐山可能不是最高的山, 然後一步一步的
往山下走, 直到最低點(local min), 所以對於一般的函數來說, 由於
可能有另外一座高山,導致也得到不同的local min2, 因此這這個演算法
並沒有得到全局最min 解. (我們可以控制走得步伐)


所以, 有了cost function 後, 我們要對他求min, 利用Gradient Descent 演算法
來求,

在若在2變數 θ0, θ1 之下, 用軟體畫出所有可能的Y=h(x), 會是一個3D
圖, 而剛好是個凸邊型, 也就是弓型, 而雖然 Gradient Descent 會求出
local min 值, 但是因為這個 cost function 是凸邊型, 所以他的local min
就是全局最小, 所以就得到我們要的答案了.


事實上 梯度下降法能用在所有地方, 不只是線性回歸.
後面還有正規方程 normal equation 也能解最小值.







[Machine Learning][Study Note][Week1] Supervised / Unsupervised Learning

Supervised Learning
=由一些已經知道解答的數據群中, 訓練出模型的方式

這主要分兩大類 1.回歸 2.分群

1.回歸=用於連續值得輸出預測

譬如你有一堆這個區域房子成交價
X= 坪數
Y=售價

你用這些歷史數據訓練出一個回歸模型, 於是你可以利用這個回歸模型去預測新的X

譬如你朋友想要在這區域賣房子, 於是他給你他房子的坪數, 你套入那個練出來的模型

得到Y, 則可以建議他以Y價格售出



2. 分群=用於離散值得輸出預測

譬如

有一堆有腫瘤病患的資料, 而假設我們X只看一個feature 叫做腫瘤直徑
(實際上會有很多feature)

而 Y就是 0=不是惡性腫瘤 1=是惡性腫瘤

所以把歷史資料都畫出來之後, 訓練出一個演算法可以分群, 以便之後出入新的病人的
腫瘤直徑, 可以預測是否為惡性


又

X=腫瘤直徑 , Y=年齡

一樣可以訓練出分群模型








Unsupervised Learning
=這是說你不知道這推數據姿料理有怎樣的規則, 你想要用演算法看看能不能找出這些資料中存在某種結構(可能這些data能分成不同群)

譬如說
每天看到的新聞網頁, 他們去其他網路上收集新聞之後, 他會被自動分到同一類, 譬如某些新聞會被分成運動, 財金 , 這就是跑了演算法之後, 發現能被分類


又或者你有一大堆的健身客戶資料, 你不知道客戶之間有甚關係, 去跑之後, 發現
這些客戶之中, 又有很多膝蓋受過傷的細分市場客戶

簡單來說,
我手上有一堆數據, 我根本不知道裡面有甚東西, 我沒給演算法答案,
請問你(演算法)你可以幫我找出這些數據當中的類型嗎?
我只告訴你演算法, 沒告訴你這些資料的答案(譬如這個人該分到哪群)