Author: maogautam

0

User defined functions(udf) in spark

UDFs or user defined functions are a simple way of adding a function into the SparkSQL language. This function operates on distributed DataFrames and works row by row....

Redshift Database connection in spark 0

Redshift Database connection in spark

This blog primarily focus on how to connect to redshift from Spark. Redshift: Amazon Redshift is a fully managed petabyte-scale data warehouse service. Redshift is designed for analytic...

hadoop logo 0

Useful commands for hadoop developer

This post combines most frequently used command for spark, emr, yarn and AWS by hadoop developer. Kill Spark  job: This command will kill all the running spark jobs.

...

0

Most common issues faced by spark developer and it’s solution

Most common issues faced by spark developer and it’s solution Timeout waiting for connection from pool Caused by: com.amazon.ws.emr.hadoop.fs.shaded.org.apache.http.conn.ConnectionPoolTimeoutException: Timeout waiting for connection from pool To resolve this...

scala-logo 0

File Operation in scala

File operation is important operation in an application. We might have to provide some configuration information or some input for application in such scenario we have to perform...

Missing Imputation in scala 0

Missing Imputation in scala

Missing imputation algorithm Read the data Get all columns name and the type of columns Replace all missing value(NA, N.A., N.A//,” ”) by null Set Boolean value for...

Missing Imputation in python 0

Missing Imputation in python

Missing imputation algorithm Read the data Get all columns name and the type of columns Replace all missing value(NA, N.A., N.A//,” ”) by null Set Boolean value for...

Spark SQL Using Parquet 0

Spark SQL Using Parquet

Today, I’m focusing on how to use parquet format in spark.  Please get the more insight about parquet format If you are new to this format. Parquet: Apache Parquet is a...