2021/04/19

ggplot2 boxplot to ggplotly

   Package "plotly" provide powerful web interactive tools in chart. ggplot2 is very intuitively for creating charts. Use ggplotly to transform ggplot2 object into plotly object is very convenient. But sometimes it's will make you surprised

For example

```{r}
library(ggplot2)
library(plotly)

df <- diamonds
p <- ggplot(df, aes_string(x = 'carat', y = 'cut')) +
  geom_boxplot()
print(p)

```




But the chart will not as you desire if you use ggplotly

 
```{r}
library(ggplot2)
library(plotly)
df <- diamonds
p <- ggplot(df, aes_string(x = 'carat', y = 'cut')) +
  geom_boxplot()
print(ggplotly(p))
```

 


plotly default set boxplot as verticle, so we use coord_flip to plot horizontal boxplot

```{r}
p <- ggplot(df, aes_string(y = 'carat', x = 'cut')) +
  geom_boxplot()+
  coord_flip()
print(ggplotly(p))

```

for grouped boxplot

```{r}
p <- ggplot(df, aes_string(y = 'carat', x = 'cut', color = 'clarity')) +
  geom_boxplot() +
  coord_flip()
print(p)
```


 

use ggplotly get stacked boxplot

```{r}
p <- ggplot(df, aes_string(y = 'carat', x = 'cut', color = 'clarity')) +
  geom_boxplot() +
  coord_flip()
print(ggplotly(p))
```

fine tune ploty layout:

```{r}
p <- ggplot(df, aes_string(y = 'carat', x = 'cut', color = 'clarity')) +
  geom_boxplot() +
  coord_flip()
print(ggplotly(p) %>% layout(boxmode='group'))
```




 

2016/12/27

Color Ramping using linear approximation

R colorRAmpPalette method with linear approximation: Input C1=(r1,g1,b1), C2=(r2,g2,b2), C3=(r3,g3,b3) and k as number of colors to generate Idea: A walk from C1 to C3 passing through C2 with equal steps. Implement The result color list denote by RC Step1 : the first half colors list distance for one step from C1 to C2 denoted by D1: (C2-C1)/((k-1)/2) RC = C1 + 1:int(k/2) * D1 Step 2: if k is odd, add C2 into this list RC = RC + C2 Step 3: the other half color list distance for one step from C2 to C3 denoted by D2: (C3-C2)/((k-1)/2) RC = RC + (C2 + 1:int(k/2) * D2) Step 4: All the element in RC must between 0 and 255

2015/08/19

updateusr method for better barplot with lines/points

When using barplot with add y2 axis for lines or points by par(new=TRUE)  as show in the example code usually get shifted.

The updateusr method from Greg Snow's TeachingDemos package is a great solution for this problem. The idea behind updateusr is define new plot region according to the ratio of one unit in each axis.

The code:


Example code:(add space 8/20)


Result:

2014/12/08

Error: Package 'XXX' was build before 3.x.x: please re-install it

For install new package which previous install use different R version cause the error:
"Error: Package 'XXX' was build before 3.x.x: please re-install it"


solution:
update.packages(checkBuilt = TRUE, ask = FALSE)

for user privilege issue:
use SUDO R, then update.packages(checkBuilt = TRUE, ask = FALSE)

Reference:
https://stackoverflow.com/questions/16987948/causes-of-error-package-was-built-before-3-0-0-please-re-install-it

2013/08/13

wafer map pattern recognization by fingerprint matching/DNA sequence matching method?

寫下來才不會又忘了 wafer map --> whole fingerprint/total DNA sequence pattern --> partial fingerprint/partial DNA sequence Q1: how to define pattern character as fingerprint's character(minutia/ridge points...) encoding wafer map pattern as DNA sequence Q2: which algorithm to apply? Reference: BFAST Fingerprint Recognition Algorithms for Partial and Full Fingerprints

2012/08/12

Boxplot in R

boxplot.stats for calculating.
the argument coef for multuplier of the IQR. 0 for extream data points in data

returning statistics are lower whisker, P25, median,P75, upper whisker

 

2012/05/31

Wave Chart

The bar plot always start at line 0, therefore use xlim for located the barplot at different x-axis position.
The code:

TO DO: add lines for input statistics

Reference : Quantile regression Chart by Philippe Grosjean, http://addictedtor.free.fr/graphiques/RGraphGallery.php?graph=109

2012/02/06

Learn something new

現在正在上Standford 的線上課程.
http://www.ml-class.org/

講的滿清楚的
學ml還可以練習英文
一舉兩得啊

2010/12/02

Defect/Failure Pattern Yield impact model

From the US patent : Patent number: 6367040
System and method for determining yield impact for semiconductor devices

It's estimate the yield impact by two step:
Step 1: calculate each defect's kill probability(the chance for die fail if this die own the defec)
Step 2: calculate the total yield loss according to the above kill probability


We can use the patent's idea for defect or pattern type's yield impact modeling.

First, use logistic regression to get the kill probability for each defect/pattern
Then, use the above kill probability to estimate the yield impact.

For defect type data, we have following data:
Each die's final result(pass or fail), which defect(s) fall in this die.
Summary data format as following:
Die's Result/Defect1/Defect2/.../DefectN
0 0 1 .... 0
1 0 1 1
...

where 0 for pass die or no defect
1 for fail die or own that defect

First we fit the logistic regression with response variable as die's result, explain variables as defect1 ... defectN
Then each defect's kill probability equals the proportion of the odds

For each fail die, we assign each defect to response the yield loss by it's kill probability.
for example, if we have 5 defect types, defect1,..., defect5 with kill probability (0.1,0.1,0.3,0.2,0.3)
if die1 is fail and with defect1 and defect5 located,
then we assign defect1 has yield loss 0.1/0.4 and defect5 has yield loss 0.3/0.4 for this die.
After summary all the fail dies, we can get the estimated yield for each defect type.

For pattern type data, replace 0 or 1 with the true yield loss, and we still have the same result.

The example R core as following

2010/09/05

Patches for speeding up R

http://www.cs.toronto.edu/~radford/R-mods.html

先down下來 找時間研究一下

2010/02/25

Tips for enhance R code

After reading Writing Efficient Programs in R and R Code optimization and Packages Creation, 3 tips by now.
1. avoid data frame (from 2nd)

2. ifelse is slower than if () { } else { } (from 1st)

3. aovid using rbind, cbind in loops, predifine a NA array or matrix(from 2nd)


Examples:

2009/11/09

EDA 常用的統計方法

先整理一下 有空再詳細寫
1.基本統計量 mean/std/min/max/percentile/
2.distribution verify and display
3.testing hypotheses
4.DOE
5.regression
6.SPC

進階分析
1.MVA(PCA/FDA)
2.Data mining

問題分類
1.Process control
2.Root cause analysis
3.Yield impact
4.Yield prediction
5.Wafer map analysis

2009/07/02

R package

The advantage of using R packages. How to create your own R packages and a little about S3 and S4 class.

http://epub.ub.uni-muenchen.de/6175/1/tr036.pdf

2009/06/16

我的31歲

30歲那年的年尾 離開了生平第一個工作的公司
31歲那年 為自己的理想打拼 為自己的人生努力

那年 讓我學看到很多 學到很多
沒有當時的歷練 就不會有現在的自己

現在雖然離開了當時努力的公司 仍是感謝有過那段經驗

2009/04/01

How many software packages is too much?

看了這篇
How many software packages is too much後的想法 一套DM package就夠了
重要的是對DM method的瞭解 知道什麼方法可以用在什麼樣的問題以及它的限制 用不同software的不同方法會有不同結果通常是因為一些預設值不同造成的 所以在用DM解決問題時 還是要回到最基本的功夫上

2009/03/19

mixed type sort

It's a common problem in semiconductor EDA analysis which plot by lot+wafer number as axis ticks. But the usual sort cannot get the right order for character embed numbers.
The package gtools has the "mixedsort" method for solving this problem.

2009/02/19

Data Modeling

好文
如何建構與評估經濟理論模型

"1.經濟理論模型必須要有實證動機;2. 經濟理論模型必須能夠透過估計檢定或是模擬校準(calibra-
tion) 等方式, 評估該理論模型的良莠"


同樣的對於統計模型 DM modeling 也是吧

2009/02/18

Axis scaling algorithm

A good look axes scale usually increase by 10^(x) or 0.5*10^(x) according to the data's range. Dorothy E. Pugh provide an algorithm (and SAS Macro) for generated suitable axis scale, I modified the SAS code into R code


reference:Dorothy E. Pugh(SUGI25), "A Robust Generalized Axis-scaling Macro"

2008/10/29

Exercise of HSAUR Chapter 1

The book: A Handbook of Statistical Analyses Using R

rle and inverse.rle

rle return a list of run length of the vector, length for the run length, value for it's corresponding value
inverse.rle rollback the original vector

example:
> x <- c(1,0,0,1,1,1,1,1,0,0,0) > (tmp <- rle(x)) Run Length Encoding lengths: int [1:4] 1 2 5 3 values : num [1:4] 1 0 1 0

>inverse.rle(tmp)
[1] 1 0 0 1 1 1 1 1 0 0 0


Application: Finding peaks in vector
getPeaks <- function(x) { tmp <- rle(diff(x)>0) change.index <- cumsum(c(1, tmp$lengths)) ind <- change.index[which(tmp$values ==FALSE)] return(ind) }

> x<-c(rnorm(300),rnorm(700,5,2)) > d1<- density(x) > ind <- getPeaks(d1$y) > plot(d1,xlab="",ylab="") > ind [1] 149 279 > points(d1$x[ind], d1$y[ind], pch = 16, col="red")

CC Copyright

創用 CC 授權條款
本著作由Chunhung Chou製作,以創用CC 姓名標示-相同方式分享 3.0 Unported 授權條款釋出。