Monday, August 13, 2018

Scikit-learn 1. Hands-on python scikit-learn: intro, using Decision Tree Regression model, MAE, overfitting, cross-validation, underfitting.

Scikit-learn is Python machine-learning library

To install it:
pip install scipy
pip install sklearn

To use Scikit-learn, we must go through several steps:

  1. Prepare data - choose appropriate data to use in model and predisction 
  2. Define - choose appropriate model (decision tree, random forest etc.)
  3. Fit - capture patterns from provided data 
  4. Predict 
  5. Evaluate - make decisions on how are predictions accurate

Our data will be:
[admin@localhost ~]$ cat > test.csv
Rooms,Price,Floors,Area
1,300,1,30
1,400,1,50
3,400,1,65
2,200,1,45
5,700,3,120
,400,2,70
,300,1,40
4,,2,95

Prepare Data

To learn how to use pandas, go to it-tuff.blogspot.com/pandas-1

python
>>> test_file_path = "~/test.csv"
>>> import pandas as pd
>>> test_data = pd.read_csv(test_file_path)
>>> # our data is having NaN values, the simpleat approach is to remove all rows with NaN data
>>> test_data = test_data.dropna(axis=0)
>>> # we need to select prediction target, by convention called "y"
>>> test_data.columns.values
>>> y = test_data.Price
>>> # we need to choose "prediction input" - features -  columns (except prediction target) which will be inputted in our model and used to make predictions
>>> # by convention features called "X"
>>> test_data_features = ['Rooms','Floors','Area']
>>> X = test_data[test_data_features]
>>> # verify data (may be something weird is out there
>>> X.head()
>>> X.describe()
>>> X.info()

Define

Our prediction target is price, price can be theoretically any real number, so we'll use Decision Tree Regression model (to read more about Decision Tree, go to - it-tuff.blogspot.com/machine-learning-1).

>>> from sklearn.tree import DecisionTreeRegressor
>>> # make our test_model to be of DecisionTreeRegressor class
>>> # DT heuristic algorithm makes optimal random decisions on each leaf (node), so result can be different for every algorithm iteration, to achieve the same result on all iterations, random_state seed must be used
>>> test_model = DecisionTreeRegressor(random_state=1)

Fit

>>> # Build a decision tree regressor from the training set (X,y) - in this step we make our Decision Tree to find patterns in the training set
>>> test_model.fit(X,y)

Predict

>>> First we'll make prediction for our training set/data to check how good model is
>>> # Making prediction for the features
>>> X
>>> # Real values are
>>> y
>>> # Model predictions are
>>> test_model.predict(X)
array([300., 400., 400., 200., 700.])
>>> 

Evaluate

If you want, you can view your decision tree model:

First export model in DOT format
>>> from sklearn.tree import export_graphviz
>>> export_graphviz(test_model,out_file="test_model.dot",feature_names=test_data_features)

Install Graphviz:
yum install graphviz
dot -Tpng test_model.dot -o test_model.png

Description of parameters in PNG file:
  1. samples - how many object are in a leaf and waiting for prediction (first leaf is having samples=5 because all 5 flats prices are waiting to be predicted)
  2. mse - several functions are available in order to measure quality of a split, mse is a default value - mean squared error - it is always non-negative, and values closer to zero are better.
  3. value - is predicted price
To evaluate our predictions, we can use many metrics, here we'll use MAE (Mean Absolute Error). To calculate MAE:
  1. Find Absolute Accuracy Error - absolute difference of price: 
    1. |actual_price - predicted_price| 
    2. This  is done for every actual price and prediction pares in training set
  2. Find mean of all errors (sum up absolute accuracy errors and divide by count of the errors)
>>> from sklearn.metrics import mean_absolute error
>>> y_true = y
>>> y_predicted = test_model.predict(X)
>>> mean_absolute_error(y_true,y_predicted)
0.0

This measure is called "in-sample" measure, because we used the same sample for both training and validating. It is bad because, for example all apartments with red door mats (if this parameter were in the data) in our sample are expensive ones, so this parameter "door mat color" will be considered while predicting apartment rent price, but it's incorrect (door mat color is not having any relation to the apartment rent price). 
In-sample prediction and validation will show that our model is ideal or close to be ideal. This is called overfitting - a model matches training data almost perfectly but does poorly on a new data. It is because each next decision tree split is having less and less corresponding values (apartments in our case). Leaves with a few apartments will  make very accurate predictions close to the actual values and this makes model perfect for training data and unreliable for new data. This is because all parameters in training model are considered to be perfect indicators for predicted value which is not the case.

On the contrary if we'll make only a few splits (low tree depth), our model will not catch important patterns in the data, so it performs poorly even in training data, this is called underfitting.

So to validate predictions correctly, we need to use different samples for prediction and validation. The simplest way to do that is to split data into prediction and validation parts (so called cross-validation):
>>> from sklearn.model_selection import train_test_split
>>> # this function splits sample data into training (by default 25% of sample size) and validating portions (mnemonics - this is TRAIN and TEST SPLIT)
>>> train_X, val_X = train_test_split(X, random_state=0)
>>> train_y, val_y = train_test_split(y, random_state=0)
>>> # now we'll use this split data to make training, prediction and validation
>>> test_model = DecisionTreeRegressor(random_state=1)
>>> test_data.fit(train_X,train_y)
>>> y_predicted = test_model.predict(val_X)
>>> mean_absolute_error(val_y,y_predicted)
150.0

As you see MAE for the in-sample data was 0.0 and for out-of-sample data is 150.0 In our data average price is 400, so error in new data (data not used during fitting) is about 37%

So we need to find compromise between overfitting and underfitting (lower MAE between training and validation data). To do that we can experiment with DecisionTreeRegressor max_leaf_nodes parameter (maximum number of leaves in our model):
def get_mae(max_leaf_nodes, train_X, val_X, train_y, val_y):   
  model = DecisionTreeRegressor(max_leaf_nodes=max_leaf_nodes, random_state=0)  
  model.fit(train_X, train_y)  
  preds_val = model.predict(val_X)  
  mae = mean_absolute_error(val_y, preds_val)  
  return(mae)  

for max_leaf_nodes in [2, 3, 4, 5]:  
  my_mae = get_mae(max_leaf_nodes, train_X, val_X, train_y, val_y)  
  print("max_leaf_nodes: %d \t MAE: %d" %(max_leaf_nodes, my_mae))  

Our data will show the same result for all max_leaf_nodes values because our test data set is too small but I think you understand importance of the above code (get_mae and for loop).
After finding the best value for max_leaf_nodes , train your model on data in-sample:
>>> test_model.fit(X,y)

Thursday, August 9, 2018

Machine Learning 1. Decision Tree (also Classification Tree or Regression Tree).

Decision Tree (DT) - is a decision support tool which consists of "leaves" (also called nodes) and "branches". Two branches form a "split". Each split corresponds to two branches with one leaf at the end of the branch. Each leaf is one of the possible values for that split:
As you can understand, DT uses heuristic algorithm, which means that DT algorithm uses practical method and this method is not guaranteed accurate or optimal but is sufficient to solve that problem.
Decision Tree Classification - helps to find class of an object using available characteristics of that object, i.e. we know which classes are existing and know parameters used to classify objects. For example: 
  1. to predict if incoming e-mail is spam or not, we use DT classification model, because we have multiple characteristics (object parameters) and want to learn class of the object and we have two classes - spam and not-spam. 
  2. to predict which number is on a sign (for simplicity think that each sign can have only numbers from 0 to 9) we'll use DT classification model, because we have characteristics of each object (pixel matrix with each pixel having it's own placement and color). So we'll try to predict to which class (10 classes - because we have 10 numbers from 0 to 9) our sign is belonging to.


Decision Tree Regression - helps to find parameters of an object using known object characteristics. In contrast to the classification, the parameter value is not a finite set of classes, but a set of real numbers. For example:
  1. to predict house price, we have a set of parameters, such as house size, floors count, placement etc. Our prediction is not predefined set of classes but a number. So we'll use DT regression model. 
There are several methods (algorithms) used to "draw" decision tree:
  1. C&RT (also CART - Classification and Regression Tree) :
    1. this method is used to draw only binary trees (each split has only two branches)
    2. on each iteration, for selected parameters (set):
      1. root of the DT is a set of all members of the model (all homes)
      2. select rule which is forming leaf (eg - home having one room or more than one rooms)
      3. we find right-branch (true - having rule - "having one room") and left-branch (false - not having rule - "having more than one room")
    3. iterations are done until we have only one branch in split or until given depth
    4. This algorithm is good for initial data analysis
  2. Random Forest - built forest is consisting of CART trees, training uses Bootstrap Aggregation or bagging method (a combination of learning models is put into the "bag" which increases the overall result). In statistics bootstrap - is any test or metric that relies on random sampling with replacement (an element may appear multiple times in a sample - this helps to estimate each element weight). Random forest can be used both for classification and regression problems. Logic of random forest:
    1. build many CART trees using different set of parameters for each DT
    2. choose most often predicted value
    3. Random Forest also automatically measures importance of the parameters (assigns score to the parameter), the sum of all scores equals 1. 

Wednesday, August 8, 2018

Pandas 1. Hands-on python pandas intro.

To install pandas:
pip install pandas

[admin@localhost ~]$ cat >  test.csv
Rooms,Price,Floors,Area
1,300,1,30
1,400,1,50
3,400,1,65
2,200,1,45
5,700,3,120
,400,2,70
,300,1,40
4,,2,95

>>> import pandas as pd
>>> test_file_path = "~/test.csv"
>>> test_data = pd.read_csv(test_file_path)
>>> type(test_data)
<class 'pandas.core.frame.DataFrame'>
>>> help(pd.DataFrame)
class DataFrame(pandas.core.generic.NDFrame)
 |  Two-dimensional size-mutable, potentially heterogeneous tabular data
 |  structure with labeled axes (rows and columns). Arithmetic operations
 |  align on both row and column labels. Can be thought of as a dict-like
 |  container for Series objects. The primary pandas data structure.
>>> test_data
<output omitted - NaN means - missing value (Not a Number))>
>>> test_data.describe()
<output omitted - to read about mean, std, percentile go to it-tuff.blogspot.com/math-1 >
>>>  test_data.columns
Index([u'Rooms', u'Price', u'Floors', u'Area'], dtype='object')
>>> test_data.columns.values
array(['Rooms', 'Price', 'Floors', 'Area'], dtype=object)
>>> help(test_data.dropna)
dropna(self, axis=0, how='any', thresh=None, subset=None, inplace=False) method of pandas.core.frame.DataFrame instance
    Remove missing values.
    ----------
    axis : {0 or 'index', 1 or 'columns'}, default 0
        Determine if rows or columns which contain missing values are
        removed.
        * 0, or 'index' : Drop rows which contain missing values.
        * 1, or 'columns' : Drop columns which contain missing value.
>>> test_data = test_data.dropna(axis=0)
>>> test_data
<<output omitted - rows with NaN values are removed>
>>> test_data.describe()
<output omitted>
>>> test_data.columns.values
array(['Rooms', 'Price', 'Floors', 'Area'], dtype=object)
>>> price = test_data.Price
>>> type(price)
<class 'pandas.core.series.Series'>
>>> help(pd.Series)
class Series(pandas.core.base.IndexOpsMixin, pandas.core.generic.NDFrame)
 |  One-dimensional ndarray with axis labels (including time series).
>>> price
<output omitted>
>>> price.describe()
<output omitted - only copy of the Price column is shown>
>>> test_data.columns.values
array(['Rooms', 'Price', 'Floors', 'Area'], dtype=object)
>>> test_data_features=['Rooms','Price']
>>> features = test_data[test_data_features]
>>> type(features)
<class 'pandas.core.frame.DataFrame'>
>>> features
<output omitted - only copies of Rooms and Price columns are shown>
>>> features.describe()
<output omitted>
>>> features.head(n=3)
<output omitted - only first 3 rows are shown>
>>> help(features.head)
head(self, n=5) method of pandas.core.frame.DataFrame instance
    Return the first `n` rows.
>>> >>> test_data.info()
<class 'pandas.core.frame.DataFrame'>
Int64Index: 5 entries, 0 to 4
Data columns (total 4 columns):
Rooms     5 non-null float64
Price     5 non-null float64
Floors    5 non-null int64
Area      5 non-null int64
dtypes: float64(2), int64(2)
memory usage: 200.0 bytes
>>> 





Math 1.  Mean, Sigma Notation, Standard Deviation and Variance, Percentile.

Mean

Mean is average - to find it - just sum and then divide to the count of summarized:
We want to find average days per months of the leap year:
First we find sum (add numbers):
31 + 29 + 31 + 30 + 31 + 30 + 31 + 31 + 30 + 31 + 30 + 31 = 366
We now that we added 12 numbers, now divide sum to the count of numbers:
366 / 12 = 30.5
So average months in leap year is having 30.5 days (Check: 30.5 * 12 = 366).

Find mean lowest temperature in Celsius in the 2017 year in Azerbaijan Ganja city:
Added lowest monthly temperatures:
- 2 - 1 + 2 + 7 + 12 + 17 + 20 + 19 + 15 + 10 + 4 + 0 = 103
Mean:
103 / 12 = 8.583 Celsius (Check: 8.583 * 12 = 102.996 ~ 103)

As you see we found mean of numbers which are of the same nature (days of the month in the first example and temperature in the month in the second example).

Sigma Notation

Σ this is sigma and it means - sum up what goes after sigma:

Σn - sum up all n's 
OK and where are n values, here there are:

    





This means sum n's and n's values are from n=1 to n=5, so:







Standard Deviation

Standard Deviation (STD) - is a measure of spread between numbers (how far are our numbers from the mean). 
To find STD (eg we have 5 flats in our building and have number of humans living in each flat (7,3,1,5,6) and we want to find STD):
  1. find mean: (7 + 3 + 1 + 5 + 6) / 5 = 4.4
  2. find differences: for each number - subtract the mean. This shows how far is this number from the mean and also shows if number lower or higher than mean:
    • 7 - 4.4 = 2.6
    • 3 - 4.4 =  -1.4
    • 1 - 4.4 = -3.4
    • 5 - 4.4 = 0.6
    • 6 - 4.4 = 1.6
  3. find squared differences: square each difference. Without this step the same negative and positive values (if any) will cancel each other and overall measure will be wrong:
    • 2.6 * 2.6 = 6.76
    • -1.4 * -1.4 = 1.96
    • -3.4 * -3.4 = 11.56
    • 0.6 * 0.6 = 0.36
    • 1.6 * 1.6 = 2.56
  4. find variance (mean of the squared differences):
    1. (6.76 + 1.96 + 11.56 + 0.36 + 2.56)/5 = 4.64
  5. find standard deviation: square root of variance:
    1. √4.64 =2.154065923 ~ 2.1541
  6. STD gives us a measure to think which number is normal (is between mean+STD & mean-STD), which is low (lower than mean-STD) or high (higher than mean+STD):
    • mean + STD = 4.4 + 2.1541 = 6.5541
    • mean - STD =  4.4 - 2.1541  = 2.2459
    • 2.2459 < 6.5541 < 7   => 7 is higher than normal for that building
    • 2.2459 < 3 < 6.5541   => 3 is normal for that building
    • 1 < 2.2459 < 6.5541   => 1 is lower than normal for that building
    • 2.2459 < 5 < 6.5541   => 5 is normal for that building
    • 2.2459 < 6 < 6.5541   => 6 is normal for that building

Percentile

Percentile - indicating the value below which a given percentage of data falls. Data itself is ordered form lower to the higher. So 95th percentile for men height is 187 cm (statistical measure), this means that 95% of men is lower than 187 cm and 5% of men is higher than 187 cm.

To find percentiles and corresponding values using nearest-rank method:

  1. order list of values, eg having list of number of humans living in 5 flats (7,3,1,5,6) : 
    • Ordered list: 1, 3, 5, 6, 7
    • Number of values N = 5
  2. find minimum, it will be 1st percentile: 1st is 1
  3. find maximum, it will be 100th percentile: 100th is 7
  4. to find  n-th percentile: n-th / 100 * N and then if not integer, round to the first higher number:
    • 25th = 25 / 100 * 5 = 1.25 ~ 2 
      • means 2nd number in list
      • so 25th percentile value is 3
    • 50th = 50 / 100 * 5 = 2.5 ~ 3 
      • means 3rd number in list
      • so 50th percentile value is 5
    • 75th = 75 / 100 * 5 = 3.75 ~ 4
      • means 4th number in list
      • so 75th percentile value is 6



















Tuesday, July 31, 2018

Python 1. Lambda, list and dictionary comprehensions.

Lambda 

Lambda is a small anonymous function which can accept any number of arguments but can only have one expression:

>>> g = lambda a,b,c :  a**b - c # literally this means lambda accepts a, b and c variables, and returns "a**b-c"
>>> g
<function <lambda> at 0x7fe8640b5410>
>>> g(1,2,3) # 1**2 - 3 = 1 - 3 = -2
-2
>>>

List comprehension

>>> squares = [n**2 for n in range(10)] 
>>> squares
[0, 1, 4, 9, 16, 25, 36, 49, 64, 81]
>>> squares = [ # more readable form of list comprehensions
... n**2 # like SQL SELECT
... for n in range(10) # like SQL FROM
... ]
>>> squares
[0, 1, 4, 9, 16, 25, 36, 49, 64, 81]
>>> squares = [
... n**2 # SELECT
... for n in range(10) # FROM
... if n%2 == 0 # WHERE
... ]
>>> squares
[0, 4, 16, 36, 64]
>>> type(squares)
<type 'list'>
>>> letters = [letter for idx,letter in enumerate("ABCDEFGHIJKLMNOPQRSTUVWXYZ")]
>>> letters
['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J', 'K', 'L', 'M', 'N', 'O', 'P', 'Q', 'R', 'S', 'T', 'U', 'V', 'W', 'X', 'Y', 'Z']
>>> 

Dictionary comprehension

>>> test_dict = {index+1 : letter for index, letter in enumerate(letters)}
>>> test_dict
{1: 'A', 2: 'B', 3: 'C', 4: 'D', 5: 'E', 6: 'F', 7: 'G', 8: 'H', 9: 'I', 10: 'J', 11: 'K', 12: 'L', 13: 'M', 14: 'N', 15: 'O', 16: 'P', 17: 'Q', 18: 'R', 19: 'S', 20: 'T', 21: 'U', 22: 'V', 23: 'W', 24: 'X', 25: 'Y', 26: 'Z'}
>>> 

Friday, July 20, 2018

Linux 1. Using Linux screen utility.

Screen is utility allowing you to open several terminal instances inside a single terminal window connection.

To install screen:
yum install screen -y

Using screen

  1. Opening new screen session: 
    1. Create screen with default name (screen will be named <pid>.<tty>.<host>):
      1. screen
    2. Create screen with custom name (<pid>.<custom-name>), i.e. wget-download. This gives ability to distinguish between present screens by name:
      1. screen -S wget-download
  2. To view all screen options:
    1. hit and release Ctrl+A and then hit ?
  3. To detach (disconnect) from current screen (you'll see "[detached from ..]" message):
    1. hit and release Ctrl+A and then hit d
  4. To list all available screens (number left to the ".pts" is screen id). (Detached) means nobody connected, (Attached) means that somebody is currently in that screen:
    1. screen -ls
  5. To reattach (reconnect) to the needed screen:
    1. By id:
      1. screen -r 12215
    2. By custom-name:
      1. screen -r wget-download
    3. Connect to already attached screen:
      1. screen -d -r 12215
  6. To lock current screen (password of the local user will be needed):
    1. hit and release Ctrl+A and then hit x
  7. To work with nested screen (screen id remains the same but you can switch between nested screens and prompt will show: screen 0 / screen 1 etc. when switching):
    1. To create nested screen:
      1. being inside screen hit and release Ctrl+A and then hit c
    2. To switch between nested screens:
      1. being inside screen hit and release Ctrl+A and then hit n (for next nested screen) or p (for previous nested screen)
    3. To list all nested screens:
      1. being inside screen hit and release Ctrl+A and then hit " (double quote - Shift+single quote)
  8. To "kill" screen:
    1. to terminate current screen type exit
    2. to terminate any screen using it's id (scree id is system pid): 
      1. kill pid

Tuesday, July 3, 2018

ASA 1. Active/Standby Failover.

1. Small FAQ

ASA Services Module is not considered in this blog post.
Failover can be Active/Active or Active/Standby. Active/Active failover must be setup in multi-context mode (per security context) and doesn't support VPN failover. Active/Standby failover supports VPN failover but all traffic goes only through ASA in active role (load-balancing is not supported). Primary and secondary units doesn't change their types (primary or secondary), only their state/role can change (i.e. secondary unit can be in active state/role due to primary unit fail).
Units have one dedicated physical port to be used as failover control link, this links must be interconnected (back-to-back without an intermediate switch). Failover control link is used for:
  1. initial failover peer discovery and negotiation
  2. replication of the configuration from active to the standby peer
  3. unit health monitoring
Both Active/Active and Active/Standby failover can be configured in stateless (no connections states are tracked) or stateful (packets and connections states are tracked and connections are not dropped when failover is done) manner. By default failover operates in stateless manner. To support stateful failover - Stateful Link must be setup.
Active unit accepts configuration changes and places the same commands to the standby unit, no configuration changes must be performed on a standby unit. If stateful failover is configured - active ASA monitors, builds and tears down all connections. Also this info also tracked and synchronized:
  1. stateful table for UDP and TCP connections
  2. ARP table and MAC mapping table
  3. routing table
  4. certain application inspection data
  5. most VPN data structures (only some client-less VPN info remains stateless)
When Active/Standby failover is used - for a  switchover to occur automatically - the active unit must become less operational than standby unit, at least one of following must occur:

  1. one of the internal (monitored) interfaces goes down
  2. an interface expansions slot fails
  3. an IPS, CSC or CX application module fails

Health messages by default are exchanged in 1 second interval. If failover control link fails, failover becomes disabled. By default, a switchover occurs when at least one interface on the active unit or within an active failover group fails.

Failover provides very effective first-hop redundancy capabilities by allowing the MAC and IP address pair on each data interface to move between the failover peers based on which unit is active at any given time. Because all physical interface connections and their configurations are identical between the members of a failover pair, active ASA unit switchovers are completely transparent to the adjacent network devices and endpoints. When you enable failover, the IP address configured on each data interface becomes the active one. When the active unit fails, the standby peer automatically assumes ownership of these addresses upon taking over the active role and seamlessly picks up transit traffic processing.

In Active/Standby failover secondary unit in active state remains active even if primary unit becomes operationally healthy. The primary unit takes over an active role only if secondary unit becomes unhealthy or if switchover is done manually.

All configuration changes must be done on the active unit (either in Primary or Secondary state).

2. Preparation

  1. When grouping two devices in failover, the following hardware parameters must be identical:
    1. exact model number
    2. number and type of physical interfaces must be the same (also expansion modules must be the same if any)
    3. all cables must be connected appropriate to the Layer 2 on both units for unit health monitoring to be held properly
    4. all hardware or software modules and software must be the same on both units
    5. amount of RAM and system flash must be the same on both units
    6. both failover peers should run the same software image during normal operation (different images are supported during upgrade) 
    7. prior to ASA8.3(1) licence features on both units must to be the same
    8. Cisco ASA 5505, ASA 5510, and ASA 5512-X appliances must have the Security Plus license installed.
    9. The state of the Encryption-3DES-AES license must match between the units. In other words, it must be either disabled or enabled on both failover peers.
  2. Choose roles for each ASA- one ASA will be primary and the other - secondary (i.e. old ASA - primary / new ASA - secondary). 
  3. Dedicate one physical interface (the same, i.e. Gi0/3 on both) on each unit for the failover control link and connect them back-to-back without an intermediate switch.  
  4. If you plan to use stateful failover - dedicate another physical interface to be used as stateful link
  5. Choose IP addresses for the primary and secondary units, used failover subnet cannot overlap with any data interfaces (one subnet per failover control and failover state links)
  6. Choose security key to encrypt failover traffic

3. Setup

3.1 Setup with separate physical interface for Failover Link and State Link

Start failover configuration on the primary node (also consider maintenance window as interface will go down while transiting to the failover active state - it takes roughly 1 minute to go into active state):
interface GigabitEthernet0/2
 no shutdown
interface GigabitEthernet0/3
 no shutdown
failover lan unit primary 
failover lan interface FailoverControl GigabitEthernet0/2 
failover link FailoverState GigabitEthernet0/3 
failover interface ip FailoverControl 172.20.0.1 255.255.255.0 standby 172.20.0.2 
failover interface ip FailoverState 172.20.1.1 255.255.255.0 standby 172.20.1.2 
failover ipsec pre-shared-key *****
failover 
     No Active mate detected
show failover | grep host
     This host: Primary - Active
     Other host: Secondary - Not Detected

Then configure failover on standby unit:
interface GigabitEthernet0/2
 no shutdown
interface GigabitEthernet0/3
 no shutdown
failover lan unit secondary 
failover lan interface FailoverControl GigabitEthernet0/2 
failover replication http 
failover link FailoverState GigabitEthernet0/3 
failover interface ip FailoverControl 172.20.0.1 255.255.255.0 standby 172.20.0.2 
failover interface ip FailoverState 172.20.1.1 255.255.255.0 standby 172.20.1.2 
failover ipsec pre-shared-key *****
failover 
     Detected an Active mate
     Beginning configuration replication from mate. 
     End configuration replication from mate.
show failover | grep host
     This host: Secondary - Standby Ready
     Other host: Other host: Primary - Active

The failover key command enables password failover encryption. Use either a string of letters, numbers, and punctuation with 1 to 63 characters or a hexadecimal value of up to 32 digits. Only use this option when running Cisco ASA Software versions earlier than 9.1(2) or deploying stateless failover.
IPSec site-to-site tunnel is more secure approach to failover link protection, so always use it in Cisco ASA Software version 9.1(2) and later. The failover ipsec pre-shared-key command enables this method of failover encryption. You must deploy stateful failover to use this feature. When using IPSec as encryption method - this tunnel is not counted in ASA maximum supported VPN count.

3.2 Setup with 1 redundant interface for both Failover Link and State Link

If you want to use redundant interface:
On primary unit:
interface GigabitEthernet0/2
 no shutdown
interface GigabitEthernet0/3
 no shutdown
interface Redundant 1
  member-interface GigabitEthernet 0/2
  INFO: security-level and IP address are cleared on GigabitEthernet0/2
  member-interface GigabitEthernet 0/3
  INFO: security-level and IP address are cleared on GigabitEthernet0/3
failover lan unit primary 
failover lan interface FailoverLink Redundant1
INFO: Non-failover interface config is cleared on Redundant1 and its sub-interfaces
failover interface ip FailoverLink 172.20.0.1 255.255.255.0 standby 172.20.0.2
failover link FailoverLink
failover ipsec pre-shared-key *****
failover 
     No Active mate detected
show failover | grep host
     This host: Primary - Active
     Other host: Secondary - Not Detected

Then configure failover on standby unit:
interface GigabitEthernet0/2
 no shutdown
interface GigabitEthernet0/3
 no shutdown
interface Redundant 1
  member-interface GigabitEthernet 0/2
  INFO: security-level and IP address are cleared on GigabitEthernet0/2
  member-interface GigabitEthernet 0/3
  INFO: security-level and IP address are cleared on GigabitEthernet0/3
failover lan unit secondary 
failover lan interface FailoverLink Redundant1
INFO: Non-failover interface config is cleared on Redundant1 and its sub-interfaces
failover interface ip FailoverLink 172.20.0.1 255.255.255.0 standby 172.20.0.2
failover link FailoverLink
failover ipsec pre-shared-key *****
failover 
     Detected an Active mate
     Beginning configuration replication from mate.
     End configuration replication from mate.
show failover | grep host
     This host: Secondary - Standby Ready
     Other host: Other host: Primary - Active

3.3 Disabling failover monitoring for interface

You can have an interface which can't be replicated (i.e. like fiber optic coming directly from an ISP). This interfaces must be unplugged from failed ASA and then plugged into currently active ASA. To exclude interface from the failover monitoring:
asa (config)# no monitor interface interface_name_here
By default, monitoring physical interfaces is enabled and monitoring subinterfaces is disabled. You can check this via: sh run all | grep monitor-interface

If you can replicate your interface (ex. Gi0/0), then on primary node:
# conf t
# int gi0/0
# ip address 192.168.0.1 255.255.255.0 standby 192.168.0.2
To check:
find name of the gi0/0:
sh nameif | grep G.*0/0
GigabitEthernet0/0       inside                   100
Check this name in failover:
sh failover | grep inside
  Interface inside (192.168.0.1): Normal (Monitored)
  Interface inside (192.168.0.2): Normal (Monitored)

You also can use standby IP to access node in Standby state.

4. Test and operate

Use the show failover command to monitor the operational state of the failover.
You can use show failover history command to investigate failover events.

Use failover execute mate command_to_execute_remotely to execute command on the standby unit (i.e.: failover exec mate show version | grep Serial). Do not execute configuration commands on the standby unit.

Use write standby - to restore standby unit proper state after accidentally performing configuration on a standby unit, this command replaces all configuration with the copy of the configuration from the active unit.

Use the failover active command on the standby unit to transit unit to the active state.
Use the no failover active command on the currently active unit  to transit unit to the standby state.