Choice of Distance Matrices in Cluster Analysis: Defining RegionsSource: Journal of Climate:;2001:;volume( 014 ):;issue: 012::page 2790DOI: 10.1175/1520-0442(2001)014<2790:CODMIC>2.0.CO;2Publisher: American Meteorological Society
Abstract: Cluster analysis is a technique frequently used in climatology for grouping cases to define classes (synoptic types or climate regimes, for example), or for grouping stations or grid points to define regions. Cluster analysis is based on some form of distance matrix, and the most commonly used metric in the climatological field has been Euclidean distances. Arguments for the use of Euclidean distances are in some ways similar to arguments for using a covariance matrix in principal components analysis: the use of the metric is valid if all data are measured on the same scale. When using Euclidean distances for cluster analysis, however, the additional assumption is made that all the variables are uncorrelated, and this assumption is frequently ignored. Two possible methods of dealing with the correlation between the variables are considered: performing a principal components analysis before calculating Euclidean distances, and calculating Mahalanobis distances using the raw data. Under certain conditions calculating Mahalanobis distances is equivalent to calculating Euclidean distances from the principal components. It is suggested that when cluster analysis is used for defining regions, Mahalanobis distances are inappropriate, and that Euclidean distances should be calculated using the unstandardized principal component scores based on only the major principal components.
|
Collections
Show full item record
| contributor author | Mimmack, Gillian M. | |
| contributor author | Mason, Simon J. | |
| contributor author | Galpin, Jacqueline S. | |
| date accessioned | 2017-06-09T15:59:23Z | |
| date available | 2017-06-09T15:59:23Z | |
| date copyright | 2001/06/01 | |
| date issued | 2001 | |
| identifier issn | 0894-8755 | |
| identifier other | ams-5823.pdf | |
| identifier uri | http://onlinelibrary.yabesh.ir/handle/yetl/4198656 | |
| description abstract | Cluster analysis is a technique frequently used in climatology for grouping cases to define classes (synoptic types or climate regimes, for example), or for grouping stations or grid points to define regions. Cluster analysis is based on some form of distance matrix, and the most commonly used metric in the climatological field has been Euclidean distances. Arguments for the use of Euclidean distances are in some ways similar to arguments for using a covariance matrix in principal components analysis: the use of the metric is valid if all data are measured on the same scale. When using Euclidean distances for cluster analysis, however, the additional assumption is made that all the variables are uncorrelated, and this assumption is frequently ignored. Two possible methods of dealing with the correlation between the variables are considered: performing a principal components analysis before calculating Euclidean distances, and calculating Mahalanobis distances using the raw data. Under certain conditions calculating Mahalanobis distances is equivalent to calculating Euclidean distances from the principal components. It is suggested that when cluster analysis is used for defining regions, Mahalanobis distances are inappropriate, and that Euclidean distances should be calculated using the unstandardized principal component scores based on only the major principal components. | |
| publisher | American Meteorological Society | |
| title | Choice of Distance Matrices in Cluster Analysis: Defining Regions | |
| type | Journal Paper | |
| journal volume | 14 | |
| journal issue | 12 | |
| journal title | Journal of Climate | |
| identifier doi | 10.1175/1520-0442(2001)014<2790:CODMIC>2.0.CO;2 | |
| journal fristpage | 2790 | |
| journal lastpage | 2797 | |
| tree | Journal of Climate:;2001:;volume( 014 ):;issue: 012 | |
| contenttype | Fulltext |