Saturday, 14 December 2019

Turn off force UTF8 encoding in pyspark

I have a python code like below to read data from Oracle using pyspark.

tableDF = spark.read \
            .format("jdbc") \
            .option("driver", "oracle.jdbc.driver.OracleDriver") \
            .option("url", "jdbc:oracle:thin:@" + hostid + ".dev.com:1521/" + databaseinstance) \
            .option("dbtable", sqlstring) \
            .option("numPartitions", 1) \
            .option("fetchsize", fetchsize) \
            .option("user", contextname) \
            .option("password", contextname) \
            .load() \

Am able to read data successfully from Oracle into the dataframe without any issues but the data got collapsed because of the force encoding to UTF-8 by pyspark which results in some of my data have the UTF-8 replacement character such as ef bfaf efbe bfef bebd which is not recognizable properly.

So, am looking to simply read data from Oracle without any encoding to my data. Basically, I dont want any forced encoding to the data am retrieving. I simply want them in the way it was stored in Oracle.

Is there a way to turn OFF this default encoding to UTF-8 by pyspark?

Any help or guidance would be much appreciated. Thanks



from Turn off force UTF8 encoding in pyspark

No comments:

Post a Comment